Named entity recognition: framing, algorithm choice, data processing, and metrics
Company: eBay
Role: Applied Scientist
Category: Machine Learning
Difficulty: easy
Interview Round: Technical Screen
Named entity recognition (NER) comes up in an ML design discussion. For concreteness, assume the text is short e-commerce text, such as product listing titles and search queries, and that the entities of interest are attributes such as brand, product type, color, size and model name. Answer the three parts in order.
### Clarifying Questions
- Which entity types are required, and is the list fixed or expected to grow?
- How much labeled data exists, and who can label more?
- Does NER run offline over listings, online in the query path, or both, and what latency budget applies online?
- Which languages must be supported?
- How will the extracted entities be used downstream (search filters, ranking features, catalog normalization)?
### Part 1 — Explain the NER problem
Define NER as a machine learning task: its input, its output, how labels are represented, and what makes it hard in this setting.
```hint Output representation
Think about how a multi-token span becomes one label per token, and how you recover spans afterward.
```
#### What This Part Should Cover
- Input and output: a token sequence in, typed spans out
- A span tagging scheme and how spans are decoded from tags
- Difficulties specific to short, noisy text: ambiguity, little context, misspellings, unseen names
### Part 2 — Algorithms and how you decide
List the main families of NER approaches, state the strengths and weaknesses of each, and explain how you would choose among them for this setting.
```hint Decision inputs
Tie the choice to data volume, latency budget, how often entity types change, and how much accuracy the downstream use needs.
```
#### What This Part Should Cover
- Rule and dictionary approaches, feature-based sequence models, neural sequence taggers, fine-tuned transformer encoders, and LLM-based extraction
- Pros and cons of each on accuracy, data needs, latency and maintainability
- A concrete decision with its reasoning and a fallback plan
### Part 3 — Data processing and metrics
Explain how you would build and process the training data, and how you would choose evaluation metrics.
```hint Two levels of evaluation
Separate how well spans are extracted from whether the extraction improves the product that consumes it.
```
#### What This Part Should Cover
- Labeling strategy and quality control, plus cheaper sources of labels
- Preprocessing: normalization, subword label alignment, class imbalance, and leakage-free splits
- Entity-level metrics (exact versus partial match, per type, micro versus macro) and downstream metrics
### What a Strong Answer Covers
- A precise task definition with a sensible tagging scheme
- A comparison of algorithm families with honest trade-offs and a justified choice
- A data plan that addresses label quality, subword alignment, imbalance and leakage
- Metrics at both the entity level and the product level, with error analysis driving iteration
### Follow-up Questions
- A new entity type must be supported with only a handful of labeled examples. What do you do?
- Offline F1 improved, but downstream search relevance did not. How do you investigate?
- How would you handle nested or overlapping entities, such as a model name that contains a brand?
- Queries are very short and often misspelled. How does that change your model and data choices?
Overview: A machine learning question on named entity recognition for short e-commerce text such as listing titles and search queries. It asks you to define the task, compare NER algorithm families and choose one, plan labeling and data processing, and select entity-level and product-level metrics.
Read the full eBay Applied Scientist interview experience this question came from