Design a multimodal embedding service
Company: Adobe
Role: Software Engineer
Category: ML System Design
Difficulty: hard
Interview Round: Technical Screen
Design a system to compute embeddings for user‑uploaded files across modalities—documents, images, and videos—where each file size is at most x MB, and persist results to a database. Describe the ingestion API, validation and preprocessing (e.g., text chunking, image resizing, video frame sampling or clip extraction), model choices per modality, batching, GPU/accelerator scheduling, and concurrency controls. Explain how you will store embeddings and metadata (e.g., vector store vs. relational/columnar DB), support similarity search, deduplicate near‑identical content, handle retries and idempotency, and manage backfills when models are updated. Include monitoring, quality evaluation, cost controls, and privacy/security considerations.
Overview: This question evaluates a candidate's competency in ML system design for building scalable, multi‑tenant, privacy‑sensitive multimodal embedding pipelines, covering ingestion and idempotency, modality-specific preprocessing, model selection and fusion, throughput engineering, storage and versioning, data hygiene, and operational concerns.
Read the full Adobe Software Engineer interview experience this question came from
Community answers
Answer by suresh.515
**# 1. Overview & Assumptions
A production service that computes embeddings for user-uploaded files across modalities — documents, images, and videos — and persists the results for similarity search and analytics. Processing is asynchronous with eventual consistency: a submission is acknowledged immediately, and embeddings become queryable once ingestion completes.
Working assumptions: each file is at most X MB (a configurable limit enforced at the API); the environment is multi-tenant and privacy-sensitive (strict isolation, encryption, retention controls); and we optimize for high, bursty throughput at controlled cost. Critically, the workload does not fit on a single GPU — both the volume of items to embed and (potentially) the model itself exceed one accelerator — so the design is multi-GPU from the start (see Section 8).
Requirements
2.1 Functional
Ingestion API to submit files/URLs, choose modality + options, and poll status — with idempotency, retries, and concurrency controls.
Validation & preprocessing per modality (documents, images, videos).
Model inference per modality, producing normalized embedding vectors.
Storage of embeddings + metadata supporting similarity search and metadata filters.
Data hygiene: dedup near-identical content, write-once semantics, and backfills on model updates.
Operations: monitoring/tracing, quality evaluation, cost controls, and privacy/security.
2.2 Non-Functional
| Quality | Target / Decision |
| --- | --- |
| Consistency | Asynchron