Data Pipeline Design for Cloud and Hybrid On-Prem Environments on AWS and Azure
Company: Disney
Role: Data Engineer
Category: System Design
Difficulty: easy
Interview Round: Onsite
Design a data pipeline that moves data from a company's operational systems into an analytics platform, where analysts, dashboards and downstream applications can use it. Expect a very detailed discussion: for every stage, name the tools and packages you would use, the alternatives you considered, and the trade-offs behind each choice.
The design is then revisited twice: first for a hybrid environment in which some systems stay on-premises while the rest runs in the cloud, and then on two cloud providers, AWS and Azure, comparing the tools available on each and the options for local and cloud deployment.
### Constraints and Clarifications
- The sources, data volumes, freshness requirements and consumers are not given. Ask for them or state your assumptions, and say how your choices would change under different answers.
- "Hybrid" means that some data sources, and possibly some processing or storage, remain in on-premises data centers while the rest of the platform runs in the cloud.
- The interviewers expect concrete services, frameworks and libraries, not only boxes and arrows.
### Clarifying Questions
- What are the sources: relational databases, application events, files from partners, third-party APIs? How much does each produce per day?
- How fresh must the data be for each consumer: daily, hourly, or within minutes?
- Who consumes the output: BI dashboards, analysts and data scientists, machine learning features, operational applications?
- Does the data include personal or otherwise sensitive fields, and do any rules require some data to stay on-premises?
- Is the organization committed to one cloud provider, or should the design stay portable?
### Part 1 — Pipeline on one cloud
Design the pipeline end to end on a single cloud provider: ingestion, storage layers, transformation, orchestration, serving to consumers, data quality and failure handling. For every stage, name the tool or package you would use and defend it against at least one alternative.
```hint Batch, streaming, or both
Let each consumer's freshness requirement decide whether a source needs a stream at all, rather than making everything real time by default.
```
```hint Managed or self-run
For each stage, compare a managed cloud service with running an open-source engine yourself, in cost, control and operational burden.
```
#### What This Part Should Cover
- Ingestion per source type (database changes, events, files, APIs) into a landing zone
- Storage layers, file and table formats, partitioning and schema evolution
- Transformation and orchestration tools, each with a defended alternative
- Data quality checks, idempotent reruns, backfills and failure alerting
### Part 2 — Hybrid environment
Some source systems, and some data, must stay on-premises, while the rest of the platform runs in the cloud. Redesign the pipeline for this hybrid environment.
```hint Crossing the boundary
Decide how data moves between the data center and the cloud: which side opens the connection, over what kind of link, and what happens while that link is slow or down.
```
```hint Where each job runs
For each transformation, ask whether it must run next to the on-premises data or can run in the cloud, and what that choice does to cost and to what data leaves the building.
```
#### Clarifying Questions for this Part
- Which data must not leave the premises, and is that driven by regulation, contracts or cost?
- What network capacity exists between the data center and the cloud, and is there a dedicated private connection?
#### What This Part Should Cover
- Connectivity and secure transfer across the boundary, with buffering when the link fails
- Placement of compute and storage on each side, and rules for what data may move
- One orchestration and monitoring view across both environments
- Consistent identity, encryption and auditing on-premises and in the cloud
### Part 3 — The same design on AWS and on Azure
Design the pipeline on AWS and on Azure. Map each stage to the services and tools you would use on each provider, compare them, and explain how you would keep the design workable across local and cloud environments.
```hint Stable core, replaceable edges
Separate the parts of the design that are tied to one provider from the parts, such as data formats, processing engines and transformation code, that could move unchanged.
```
#### What This Part Should Cover
- A stage-by-stage mapping of services on each provider
- Differences that actually change the design, not only the names of services
- Portability choices and what they cost
- How local development and testing mirror the cloud deployment
### What a Strong Answer Covers
- Choices driven by requirements: freshness, volume and consumers decide batch versus streaming and the tools
- A concrete trade-off and an alternative behind every tool choice
- Correctness under reruns, late and duplicate data, schema changes and backfills
- Data quality, security and sensitive-data handling built into the pipeline rather than added at the end
- Operability across environments: orchestration, monitoring, lineage and cost control
### Follow-up Questions
- An on-premises source changes its schema without notice. How does the pipeline detect it, and what happens to the downstream tables?
- The link to the data center is down for several hours. What happens to freshness, and how does the pipeline catch up without duplicating data?
- The monthly cloud bill doubles after the migration. Where do you look first?
- How would you backfill two years of history for a newly added column without disrupting the daily runs?
Overview: A senior data engineering design question: build a data pipeline from operational sources to analytics on one cloud, then rework it for a hybrid on-premises and cloud environment and for both AWS and Azure. It tests tool selection and trade-offs, ingestion and storage design, data quality, security and portability.
Read the full Disney Data Engineer interview experience this question came from