Design a Model That Predicts Whether a Member Clicks a Displayed Job

Read the full interview experience this question came from →

Quick Overview

An ML system design question: predict whether a member will click a job displayed to them, using member features, job content and interaction history. It tests defining the supervision signal and training data, choosing a loss, preventing data leakage, building member and job embeddings, and laying out the training and serving components.

Design a Model That Predicts Whether a Member Clicks a Displayed Job

Company: LinkedIn

Role: Machine Learning Engineer

Category: ML System Design

Difficulty: hard

Interview Round: Technical Screen

A professional networking platform displays job postings to its members. Design a machine-learning system that predicts whether a member will click on a job that is displayed to them, using the member's features, the job's content and the member's interaction history. Assume the predicted click probability is used to rank the jobs the member sees. You are expected to define the problem yourself, including: - what data you need, including the historical interaction data; - the supervision signal: what makes a training example positive or negative; - the system components, offline and online; - how the model is trained, and with which loss function; - how you prevent data leakage; - how you obtain member and job embeddings. The original discussion lasted about 20 minutes, so prioritize. ```hint What is one example? A click only means something relative to a job that was actually shown. Decide which logged event creates a training row, and what you do with jobs that were displayed but never really seen. ``` ```hint Leakage hides in time For every feature, ask when its value was computed relative to the moment the job was displayed. ``` ```hint Embeddings for a job posted an hour ago Member and job vectors can come from content, from engagement, or from both. Consider which source still works for a posting that nobody has interacted with yet. ``` ### Clarifying Questions - Which surfaces display the jobs, and should each have its own model or should one model take the surface as a feature? - Is a click the real objective, or a proxy for applications and hires that matter more to members and job posters? - How many candidate jobs must be scored per page view, and within what latency? - How much interaction history does a typical member have, and what share of members and jobs are new on any given day? ### What a Strong Answer Covers - A precise supervision signal: which impressions become examples, how clicks are attributed, and how unseen impressions and position bias are handled - The data sources (member profile, job content, interaction logs, context) and how they are joined to each impression - Member, job and member-job features, plus a concrete method for member and job embeddings that handles cold start - A model choice and loss function justified by how the score is used, with calibration and class imbalance addressed - Specific leakage risks and the mechanisms that prevent them, including how data is split for training and evaluation - The offline and online components: candidate generation, feature store, scoring, logging, retraining - Offline metrics, an online experiment, and guardrails beyond click-through rate ### Follow-up Questions - The business now wants more applications rather than more clicks. What changes in the labels, the model and the evaluation? - Jobs in the top slots get clicked more because of where they appear. How do you keep the model from learning position instead of relevance? - A job posted ten minutes ago has no interactions. How does it get a useful embedding and score? - Offline AUC improves, but the online click-through rate drops after launch. What do you investigate?

Overview: An ML system design question: predict whether a member will click a job displayed to them, using member features, job content and interaction history. It tests defining the supervision signal and training data, choosing a loss, preventing data leakage, building member and job embeddings, and laying out the training and serving components.

Read the full LinkedIn Machine Learning Engineer interview experience this question came from

|Home/ML System Design/LinkedIn
LinkedIn logo
LinkedIn
Oct 1, 2026
hardMachine Learning EngineerTechnical ScreenML System Design
0
0

A professional networking platform displays job postings to its members. Design a machine-learning system that predicts whether a member will click on a job that is displayed to them, using the member's features, the job's content and the member's interaction history. Assume the predicted click probability is used to rank the jobs the member sees.

You are expected to define the problem yourself, including:

  • what data you need, including the historical interaction data;
  • the supervision signal: what makes a training example positive or negative;
  • the system components, offline and online;
  • how the model is trained, and with which loss function;
  • how you prevent data leakage;
  • how you obtain member and job embeddings.

The original discussion lasted about 20 minutes, so prioritize.

Clarifying Questions Guidance

  • Which surfaces display the jobs, and should each have its own model or should one model take the surface as a feature?
  • Is a click the real objective, or a proxy for applications and hires that matter more to members and job posters?
  • How many candidate jobs must be scored per page view, and within what latency?
  • How much interaction history does a typical member have, and what share of members and jobs are new on any given day?

What a Strong Answer Covers Guidance

  • A precise supervision signal: which impressions become examples, how clicks are attributed, and how unseen impressions and position bias are handled
  • The data sources (member profile, job content, interaction logs, context) and how they are joined to each impression
  • Member, job and member-job features, plus a concrete method for member and job embeddings that handles cold start
  • A model choice and loss function justified by how the score is used, with calibration and class imbalance addressed
  • Specific leakage risks and the mechanisms that prevent them, including how data is split for training and evaluation
  • The offline and online components: candidate generation, feature store, scoring, logging, retraining
  • Offline metrics, an online experiment, and guardrails beyond click-through rate

Follow-up Questions Guidance

  • The business now wants more applications rather than more clicks. What changes in the labels, the model and the evaluation?
  • Jobs in the top slots get clicked more because of where they appear. How do you keep the model from learning position instead of relevance?
  • A job posted ten minutes ago has no interactions. How does it get a useful embedding and score?
  • Offline AUC improves, but the online click-through rate drops after launch. What do you investigate?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...