Design a multimodal embedding service

Read the full interview experience this question came from →

Quick Overview

This question evaluates a candidate's competency in ML system design for building scalable, multi‑tenant, privacy‑sensitive multimodal embedding pipelines, covering ingestion and idempotency, modality-specific preprocessing, model selection and fusion, throughput engineering, storage and versioning, data hygiene, and operational concerns.

Design a multimodal embedding service

Company: Adobe

Role: Software Engineer

Category: ML System Design

Difficulty: hard

Interview Round: Technical Screen

Design a system to compute embeddings for user‑uploaded files across modalities—documents, images, and videos—where each file size is at most x MB, and persist results to a database. Describe the ingestion API, validation and preprocessing (e.g., text chunking, image resizing, video frame sampling or clip extraction), model choices per modality, batching, GPU/accelerator scheduling, and concurrency controls. Explain how you will store embeddings and metadata (e.g., vector store vs. relational/columnar DB), support similarity search, deduplicate near‑identical content, handle retries and idempotency, and manage backfills when models are updated. Include monitoring, quality evaluation, cost controls, and privacy/security considerations.

Overview: This question evaluates a candidate's competency in ML system design for building scalable, multi‑tenant, privacy‑sensitive multimodal embedding pipelines, covering ingestion and idempotency, modality-specific preprocessing, model selection and fusion, throughput engineering, storage and versioning, data hygiene, and operational concerns.

Read the full Adobe Software Engineer interview experience this question came from

Community answers

Answer by suresh.515

**# 1. Overview & Assumptions A production service that computes embeddings for user-uploaded files across modalities — documents, images, and videos — and persists the results for similarity search and analytics. Processing is asynchronous with eventual consistency: a submission is acknowledged immediately, and embeddings become queryable once ingestion completes. Working assumptions: each file is at most X MB (a configurable limit enforced at the API); the environment is multi-tenant and privacy-sensitive (strict isolation, encryption, retention controls); and we optimize for high, bursty throughput at controlled cost. Critically, the workload does not fit on a single GPU — both the volume of items to embed and (potentially) the model itself exceed one accelerator — so the design is multi-GPU from the start (see Section 8). Requirements 2.1 Functional Ingestion API to submit files/URLs, choose modality + options, and poll status — with idempotency, retries, and concurrency controls. Validation & preprocessing per modality (documents, images, videos). Model inference per modality, producing normalized embedding vectors. Storage of embeddings + metadata supporting similarity search and metadata filters. Data hygiene: dedup near-identical content, write-once semantics, and backfills on model updates. Operations: monitoring/tracing, quality evaluation, cost controls, and privacy/security. 2.2 Non-Functional | Quality | Target / Decision | | --- | --- | | Consistency | Asynchron
|Home/ML System Design/Adobe
Adobe logo
Adobe
Sep 6, 2025
hardSoftware EngineerTechnical ScreenML System Design
11
0

System Design: Multimodal Embedding Pipeline for Documents, Images, and Videos

You are designing a production service that computes embeddings for user‑uploaded files across modalities—documents, images, and videos—and persists results for search and analytics.

Assume:

  • Each file is at most x MB (a configurable limit enforced at the API).
  • Processing is asynchronous with eventual consistency (embeddings become available after ingestion completes).
  • Multi‑tenant, privacy‑sensitive environment.

Requirements

  1. Ingestion API
    • Endpoints to submit files or URLs, specify modality and options, and poll status.
    • Idempotency, retries, and concurrency controls.
  2. Validation and Preprocessing
    • Documents: text extraction, language detection, tokenization, chunking with overlap, boilerplate removal.
    • Images: format normalization, orientation, resizing/cropping, optional multi‑crop/tiling.
    • Videos: frame sampling or clip extraction, optional ASR for audio track, keyframe/shot detection.
  3. Model Choices per Modality
    • Text/document embedding model.
    • Image embedding model.
    • Video embedding model or aggregation of frame embeddings; optional fusion with ASR text.
    • Consider a single cross‑modal space vs. per‑modality spaces.
  4. Throughput Engineering
    • Batching, GPU/accelerator scheduling, and backpressure.
    • Worker pools for CPU preprocessing vs. GPU inference.
  5. Storage Design
    • How to store embeddings and metadata (vector store vs. relational/columnar DB).
    • Schema, versioning, and multi‑tenancy.
    • Support similarity search and metadata filters.
  6. Data Hygiene and Robustness
    • Deduplicate near‑identical content (files, chunks, frames).
    • Retries, idempotency, and exactly‑once/write‑once semantics.
    • Backfills when models are updated (index rebuilds, blue/green swaps).
  7. Operations
    • Monitoring, alerting, and tracing.
    • Quality evaluation (offline metrics, canaries, A/B tests).
    • Cost controls (batching, quantization, index compression, TTLs).
    • Privacy/security (encryption, access control, retention policies).

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...