PracHub
QuestionsLearningGuidesInterview Prep
|Home/Data Manipulation (SQL/Python)/EY

Architect cloud data ingestion patterns

Last updated: Jul 17, 2026

Quick Overview

This question evaluates a candidate's ability to design end-to-end cloud data ingestion and serving architectures, including streaming and batch patterns, CDC/event sourcing/micro-batch trade-offs, partitioning and compaction strategies, idempotency, schema evolution, PII tokenization, cost controls, and incident simulation for chaos testing.

  • medium
  • EY
  • Data Manipulation (SQL/Python)
  • Data Scientist

Architect cloud data ingestion patterns

Company: EY

Role: Data Scientist

Category: Data Manipulation (SQL/Python)

Difficulty: medium

Interview Round: Technical Screen

Propose a cloud data ingestion and serving pattern for streaming and batch on your preferred cloud (AWS/Azure/GCP). Choose between CDC, event sourcing, or micro‑batch for upstream systems, justify partitioning/compaction, and show how you ensure idempotency, schema evolution (e.g., optional fields), and PII tokenization. Include cost controls (storage tiering, TTL, file size targets) and an incident you would simulate in chaos testing.

Quick Answer: This question evaluates a candidate's ability to design end-to-end cloud data ingestion and serving architectures, including streaming and batch patterns, CDC/event sourcing/micro-batch trade-offs, partitioning and compaction strategies, idempotency, schema evolution, PII tokenization, cost controls, and incident simulation for chaos testing.

Related Interview Questions

  • Design logical model and consumption - EY (medium)
  • Map sources to functional dataset with SQL - EY (medium)
  • Design a data platform enablement - EY (medium)
|Home/Data Manipulation (SQL/Python)/EY

Architect cloud data ingestion patterns

EY logo
EY
Oct 13, 2025, 9:49 PM
mediumData ScientistTechnical ScreenData Manipulation (SQL/Python)
4
0

Propose a cloud data ingestion and serving pattern for streaming and batch on your preferred cloud (AWS/Azure/GCP). Choose between CDC, event sourcing, or micro‑batch for upstream systems, justify partitioning/compaction, and show how you ensure idempotency, schema evolution (e.g., optional fields), and PII tokenization. Include cost controls (storage tiering, TTL, file size targets) and an incident you would simulate in chaos testing.

Loading comments...

Browse More Questions

More Data Manipulation (SQL/Python)•More EY•More Data Scientist•EY Data Scientist•EY Data Manipulation (SQL/Python)•Data Scientist Data Manipulation (SQL/Python)

Write your answer

Your first approved answer each day earns 20 XP.

Sign in to write your answer.
PracHub

Master your tech interviews with 9,000+ real questions from top companies.

Product

  • Questions
  • Learning Tracks
  • Interview Guides
  • Resources
  • Premium
  • For Universities

Browse

  • By Company
  • By Role
  • By Category
  • Topic Hubs
  • SQL Questions
  • AI Coding Questions
  • Compare Platforms
  • Discord Community

Support

  • support@prachub.com
  • (916) 541-4762

Legal

  • Privacy Policy
  • Terms of Service
  • About Us

© 2026 PracHub. All rights reserved.