Implement Simple Linear Regression with Gradient Descent from Scratch

Quick Overview

Write simple linear regression trained with gradient descent from scratch, choosing the loss, learning rate and stopping rule yourself. Tests deriving the gradients correctly, writing a runnable training loop, understanding feature scaling and convergence, and verifying results against a closed-form fit.

Implement Simple Linear Regression with Gradient Descent from Scratch

Company: Databricks

Role: Machine Learning Engineer

Category: Machine Learning

Difficulty: medium

Interview Round: Onsite

Write simple linear regression trained with gradient descent, from scratch. Given training inputs `x` (one feature per example) and targets `y`, learn a slope `w` and an intercept `b` for the model $\hat{y} = w x + b$. Run gradient descent on a loss you choose and justify, and return the learned parameters. Do not call a library routine that fits the model for you. The report records only the task itself. The loss, the stopping rule and the hyperparameters were not described, so choose them and explain your choices. ```hint Start from the loss Write the loss as an average over examples, then differentiate it with respect to each parameter before writing any code. ``` ```hint Watch the scale Consider what happens with a fixed learning rate when the feature values are in the thousands rather than close to one. ``` ### Clarifying Questions - Is this one feature only (simple regression), or should the code generalize to several features? - May I use NumPy for vector arithmetic, or must it be plain Python? - Should training be full-batch, stochastic or mini-batch? - How should the result be checked: against a closed-form solution, on held-out data, or with a unit test? ### What a Strong Answer Covers - A correct derivation of the gradients for both parameters from the chosen loss - Clean, runnable code for the training loop, with a sensible initialization, learning rate and stopping rule - Why feature scaling matters for convergence, and how to map the parameters back to the original scale - A way to verify the result, such as a comparison with the closed-form least-squares answer - Edge cases: constant inputs, divergence from a learning rate that is too large, and empty or mismatched inputs - Time and memory cost per iteration ### Follow-up Questions - How does the code change for many features, and when would you use the normal equations instead? - How would you add L2 regularization, and should the intercept be regularized? - What changes with stochastic or mini-batch updates, and how would you choose the batch size?

Overview: Write simple linear regression trained with gradient descent from scratch, choosing the loss, learning rate and stopping rule yourself. Tests deriving the gradients correctly, writing a runnable training loop, understanding feature scaling and convergence, and verifying results against a closed-form fit.

|Home/Machine Learning/Databricks
Databricks logo
Databricks
Sep 11, 2026
mediumMachine Learning EngineerOnsiteMachine Learning
0
0

Write simple linear regression trained with gradient descent, from scratch. Given training inputs x (one feature per example) and targets y, learn a slope w and an intercept b for the model y^=wx+b\hat{y} = w x + b. Run gradient descent on a loss you choose and justify, and return the learned parameters. Do not call a library routine that fits the model for you.

The report records only the task itself. The loss, the stopping rule and the hyperparameters were not described, so choose them and explain your choices.

Clarifying Questions Guidance

  • Is this one feature only (simple regression), or should the code generalize to several features?
  • May I use NumPy for vector arithmetic, or must it be plain Python?
  • Should training be full-batch, stochastic or mini-batch?
  • How should the result be checked: against a closed-form solution, on held-out data, or with a unit test?

What a Strong Answer Covers Guidance

  • A correct derivation of the gradients for both parameters from the chosen loss
  • Clean, runnable code for the training loop, with a sensible initialization, learning rate and stopping rule
  • Why feature scaling matters for convergence, and how to map the parameters back to the original scale
  • A way to verify the result, such as a comparison with the closed-form least-squares answer
  • Edge cases: constant inputs, divergence from a learning rate that is too large, and empty or mismatched inputs
  • Time and memory cost per iteration

Follow-up Questions Guidance

  • How does the code change for many features, and when would you use the normal equations instead?
  • How would you add L2 regularization, and should the intercept be regularized?
  • What changes with stochastic or mini-batch updates, and how would you choose the batch size?
Loading comments...