Retrieve and Rerank Tables and Notebooks for a Data-Aware Coding Agent
Company: Databricks
Role: Machine Learning Engineer
Category: ML System Design
Difficulty: medium
Interview Round: Onsite
An internal coding agent helps employees write code and answer questions. When someone asks a data question, such as how a metric changed over some period, the agent must find the most suitable tables and notebooks in the company's data platform so it can answer. Design the retrieval system that supplies those assets to the agent.
The design centers on retrieval plus a reranker. The discussion goes into how to build the index, what information about each asset to put into it, how to handle notebooks and tables that are noisy or not actually helpful, and how to design the metrics.
### Clarifying Questions
- Roughly how many tables and notebooks exist, and how quickly do they change?
- What metadata is available per asset: schemas, column descriptions, owners, usage logs, lineage?
- Does the agent consume the top few assets directly, or does a human pick from a list?
- Must results respect the asking user's access permissions?
- Is there any historical signal of which assets answered which questions, such as the queries users actually ran afterwards?
### Part 1 — What to index, and how
Decide which information about tables and notebooks goes into the index, how each asset is represented or split, and how the index stays fresh.
```hint Questions rarely name tables
Think about what words in a user's question could possibly match inside each kind of asset, and what you would have to add to an asset so that they do.
```
#### What This Part Should Cover
- A representation for each asset type, covering both text and structured signals
- Enriching assets so that business language in questions can match technical names
- Incremental indexing and freshness
### Part 2 — Retrieval and reranking
Design the path from a question to the handful of assets the agent receives.
```hint Two vocabularies
Identifiers like snake_case column names and natural-language questions rarely share tokens. Consider more than one way to generate candidates.
```
#### What This Part Should Cover
- Candidate generation aimed at recall, from more than one source
- The reranker's inputs, the model, and where its training labels come from
- What the agent receives, and the latency budget
### Part 3 — Noisy and unhelpful assets
Many notebooks are scratch work, broken or duplicated, and many tables are stale, deprecated or test copies. How do you keep them from crowding out the right answers?
```hint Look outside the text
Look for signals of an asset's quality that exist outside the asset's own text.
```
#### What This Part Should Cover
- Quality signals, and whether each acts as a filter or as a ranking feature
- Handling near-duplicates and stale or deprecated assets
- Cases where an asset's description disagrees with its actual content
### Part 4 — Metrics
How do you measure whether the system works?
```hint Stage vs end to end
Separate what you can measure at each stage from what you can measure only end to end.
```
#### What This Part Should Cover
- Retrieval and ranking metrics, and the labels behind them
- End-to-end answer quality
- Online metrics and guardrails
### What a Strong Answer Covers
- Asset representations that bridge the gap between natural-language questions and technical metadata
- A hybrid candidate-generation stage followed by a feature-rich reranker, with a credible source of training labels
- Access control enforced before any asset reaches the agent
- Systematic handling of noisy, duplicate and stale assets using signals beyond the text
- Metrics at each stage plus end-to-end answer correctness, both offline and online
### Follow-up Questions
- How would you handle a question that needs a join across two tables that never appear together?
- How would you cold-start a newly created table that has no usage history?
- How would you use the agent's own successes and failures to improve the reranker over time?
Overview: Design the retrieval system that lets an internal coding agent find the right tables and notebooks to answer data questions. Tests what to index for each asset type, hybrid retrieval plus reranking, handling noisy or stale assets, access control, and metric design from retrieval to final answers.