Design a Kafka-Fed Event Analytics System for Long-Term Analysis

Quick Overview

A system design exercise to build an event analytics system that ingests very high-volume, concurrent event streams from Kafka and stores them for long-term analysis. It tests Kafka topics, partitions, consumers and consumer groups, delivery guarantees, and durable, queryable storage design.

Design a Kafka-Fed Event Analytics System for Long-Term Analysis

Company: Attentive

Role: Software Engineer

Category: System Design

Difficulty: hard

Interview Round: Onsite

Design an event analytics system. The system connects to Kafka streams, must ingest a very large volume of highly concurrent event data, and stores that data for long-term analysis. Expect the interviewer to probe your understanding of Kafka itself: topics, partitions, consumers and consumer groups. ### Clarifying Questions - What kinds of events arrive, and who produces them: internal services, client applications, or both? - What must the long-term analysis support: ad-hoc queries over raw events, fixed dashboards, or both? How fresh must the results be? - What peak event rate and average event size should the design handle, and how long must data be kept? - Must every event be counted exactly once, or is at-least-once delivery with deduplication acceptable? - Can events be corrected or deleted after they are stored, for example to honor a data-deletion request? ### Part 1 — Ingestion and long-term storage Design the path from the Kafka topics to durable storage and on to analysis: what consumes the streams, how the data is validated, transformed and written, where it lives for the long term, and how analysts query it. ```hint Store for the reader Think about the file format, layout and partitioning that make long-range analytical scans cheap. ``` ```hint Plan for replays Assume a consumer crashes halfway through a batch and restarts. Decide what makes reprocessing the same events safe. ``` #### What This Part Should Cover - The stages from Kafka to storage: consumers, batching, file format and partitioning - Schema management and evolution - Idempotent writes and deduplication when events are replayed - The query layer, and the trade-off between freshness and cost ### Part 2 — Kafka topics, partitions and consumer groups Now go deep on Kafka. How would you lay out topics, choose partition counts and pick message keys? How do consumers and consumer groups divide the work, what happens when a consumer joins, leaves or falls behind, and how are offsets managed? ```hint Let the key follow the ordering you need Decide which events must be processed in order relative to one another, and let that drive the partition key. ``` ```hint Find the ceiling on parallelism Work out what limits how many consumers in one group can do useful work at the same time. ``` #### What This Part Should Cover - Topic and partition design, key choice, and hot partitions - Consumer group mechanics: assignment, rebalancing and the limit on parallelism - Offset commits and the delivery guarantee they produce - Consumer lag, backpressure and replication settings ### What a Strong Answer Covers - Requirements and capacity estimates stated up front and used to size partitions and storage - Correct Kafka mechanics, not just component names - End-to-end reasoning about the delivery guarantee, from producer to stored data - Cost-aware long-term storage and retention - Failure handling and observability, including consumer lag and bad events ### Follow-up Questions - One key suddenly produces a large share of all events and one partition falls far behind. What do you do? - You need more partitions on a live topic. What breaks, and how do you manage the change? - A bug in the transformation corrupted a day of stored data. How do you reprocess it without double counting? - Analysts now also want counts that are only seconds old. How does the design change?

Overview: A system design exercise to build an event analytics system that ingests very high-volume, concurrent event streams from Kafka and stores them for long-term analysis. It tests Kafka topics, partitions, consumers and consumer groups, delivery guarantees, and durable, queryable storage design.

|Home/System Design/Attentive
Attentive logo
Attentive
Sep 19, 2026
hardSoftware EngineerOnsiteSystem Design
1
0

Design an event analytics system. The system connects to Kafka streams, must ingest a very large volume of highly concurrent event data, and stores that data for long-term analysis. Expect the interviewer to probe your understanding of Kafka itself: topics, partitions, consumers and consumer groups.

Clarifying Questions Guidance

  • What kinds of events arrive, and who produces them: internal services, client applications, or both?
  • What must the long-term analysis support: ad-hoc queries over raw events, fixed dashboards, or both? How fresh must the results be?
  • What peak event rate and average event size should the design handle, and how long must data be kept?
  • Must every event be counted exactly once, or is at-least-once delivery with deduplication acceptable?
  • Can events be corrected or deleted after they are stored, for example to honor a data-deletion request?

Part 1 — Ingestion and long-term storage

Design the path from the Kafka topics to durable storage and on to analysis: what consumes the streams, how the data is validated, transformed and written, where it lives for the long term, and how analysts query it.

What This Part Should Cover Guidance

  • The stages from Kafka to storage: consumers, batching, file format and partitioning
  • Schema management and evolution
  • Idempotent writes and deduplication when events are replayed
  • The query layer, and the trade-off between freshness and cost

Part 2 — Kafka topics, partitions and consumer groups

Now go deep on Kafka. How would you lay out topics, choose partition counts and pick message keys? How do consumers and consumer groups divide the work, what happens when a consumer joins, leaves or falls behind, and how are offsets managed?

What This Part Should Cover Guidance

  • Topic and partition design, key choice, and hot partitions
  • Consumer group mechanics: assignment, rebalancing and the limit on parallelism
  • Offset commits and the delivery guarantee they produce
  • Consumer lag, backpressure and replication settings

What a Strong Answer Covers Guidance

  • Requirements and capacity estimates stated up front and used to size partitions and storage
  • Correct Kafka mechanics, not just component names
  • End-to-end reasoning about the delivery guarantee, from producer to stored data
  • Cost-aware long-term storage and retention
  • Failure handling and observability, including consumer lag and bad events

Follow-up Questions Guidance

  • One key suddenly produces a large share of all events and one partition falls far behind. What do you do?
  • You need more partitions on a live topic. What breaks, and how do you manage the change?
  • A bug in the transformation corrupted a day of stored data. How do you reprocess it without double counting?
  • Analysts now also want counts that are only seconds old. How does the design change?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...