Debug a From-Scratch Logistic Regression: Sigmoid Derivative and Gradients
Company: LinkedIn
Role: Machine Learning Engineer
Category: Machine Learning
Difficulty: hard
Interview Round: Technical Screen
You are given a small NumPy implementation of binary logistic regression trained with full-batch gradient descent. It should learn a weight vector `w` and a bias `b` that minimize the mean log loss (binary cross-entropy) over the training set, but the script does not train correctly. Debug it: find each defect, explain why it is wrong, fix it, and show how you would confirm that the fixed version is correct.
The exact code used in the interview was not reported. Practice on the representative listing below, which has the same structure as the reported exercise: a sigmoid, its derivative, a loss, a gradient computation and a training loop.
```python
import numpy as np
def sigmoid(z):
return 1.0 / (1.0 + np.exp(-z))
def sigmoid_derivative(z):
s = sigmoid(z)
return s * (1 - z)
def compute_loss(X, y, w, b):
p = sigmoid(X @ w + b)
return -np.mean(y * np.log(p) + (1 - y) * np.log(1 - p))
def compute_gradients(X, y, w, b):
z = X @ w + b
p = sigmoid(z)
error = (p - y) * sigmoid_derivative(z)
dw = X @ error
db = np.sum(error)
return dw, db
def train(X, y, lr=0.1, epochs=1000):
n, d = X.shape
w = np.zeros(d)
b = 0.0
for _ in range(epochs):
dw, db = compute_gradients(X, y, w, b)
w += lr * dw
b += lr * db
return w, b
```
### Constraints and Clarifications
- `X` is a float64 array of shape `(n, d)`, `y` is an array of shape `(n,)` holding 0 or 1, `w` has shape `(d,)`, and `b` is a scalar.
- The objective is the unregularized mean log loss over the `n` examples.
- Use NumPy only; no automatic differentiation library.
- Assume features are not standardized, so the linear score `X @ w + b` can be large in magnitude.
### Clarifying Questions
- Is the loss meant to be the mean or the sum over examples, so that the gradient and the learning rate can be scaled consistently?
- Are labels encoded as 0 and 1, or as -1 and +1?
- Is it enough to make this script train, or should each helper also be robust enough to reuse elsewhere?
### Part 1 — The sigmoid and its derivative
Review `sigmoid` and `sigmoid_derivative`. Derive the derivative of the sigmoid by hand, fix any mistake in the code, and check how both functions behave for very large positive and very large negative inputs.
```hint Test at a known point
At `z = 0` you can compute the exact value of both functions in your head. Compare that with what the code returns, then try `z = 1000` and `z = -1000`.
```
#### What This Part Should Cover
- A correct derivation of the sigmoid's derivative
- The defect in `sigmoid_derivative` and how it shows up in values
- Overflow and saturation behavior of `sigmoid`, and a numerically stable formulation
- Small unit tests that pin both functions down
### Part 2 — The gradients and the update
Review `compute_loss`, `compute_gradients` and `train`. Derive the gradient of the mean log loss with respect to `w` and `b`, then fix whatever stops the script from running and the loss from decreasing.
```hint Simplify before you code
Derive the derivative of a single example's log loss with respect to its linear score `z`, going through the sigmoid, and simplify it completely before comparing it with `error` in the code.
```
```hint Track the shapes
Write down the shape of every array in `compute_gradients` for `n = 5` and `d = 3`. Then ask what the same code does when `n == d`.
```
```hint Watch the loss
Run a few iterations on a tiny dataset and print `compute_loss` after each one.
```
#### What This Part Should Cover
- The gradient derivation for log loss with a sigmoid output, and what the code's formula computes instead
- Correct vectorized shapes, and a gradient scaled consistently with a mean loss
- The direction of the parameter update
- A loss that stays finite when predictions saturate
### What a Strong Answer Covers
- A systematic debugging order: reproduce on a tiny input, isolate each function, test against hand-computed values, then fix
- Every defect named with the reason it is wrong, not just a rewritten script
- Numerical stability in both the sigmoid and the loss
- Verification with a finite-difference gradient check and an end-to-end training sanity check
- Prioritizing the defects that stop training within a short time box
### Follow-up Questions
- How would you add L2 regularization, and should the bias be regularized?
- How do the gradients change for multi-class softmax regression?
- The training set no longer fits in memory. What do you change in `train`?
- The gradient check passes, yet the trained model predicts the majority class for every example. What do you investigate next?
Overview: A machine learning coding exercise: debug a from-scratch NumPy logistic regression, from the sigmoid and its derivative through the gradient computation and the training update. It tests deriving the log-loss gradient, spotting shape, scale and sign errors, numerical stability, and verifying a fix with a finite-difference gradient check.
Read the full LinkedIn Machine Learning Engineer interview experience this question came from