Design a Feature Flag Service with Percentage Rollouts and Allow/Block Lists

Read the full interview experience this question came from →

Quick Overview

System design question: design a feature flag service, similar in spirit to Amazon's Weblab, that lets teams turn features on or off, roll them out to a percentage of users, and force them on or off with allow and block lists. It probes where flag checks run, how users are assigned consistently, and how changes reach every service safely.

Design a Feature Flag Service with Percentage Rollouts and Allow/Block Lists

Company: Stripe

Role: Software Engineer

Category: System Design

Difficulty: hard

Interview Round: Technical Screen

Design a feature flag service, similar in spirit to Amazon's Weblab, for a company that runs many backend services. Engineers create and edit flags in a management console. For each flag they can turn the feature on or off for everyone, roll it out to a percentage of users, and force it on or off for specific users with an allow list and a block list. Application services check flags while they handle requests. ```hint Start from the access pattern Compare how often a flag changes with how often it is checked, and how stale a check may be. Let those numbers, not a design you have used for a different problem, decide where a check should be computed. ``` ```hint Same user, same answer A user who gets the feature on one request should get it on the next, on any server, and should keep it when the rollout grows. Think about how to guarantee that without storing an assignment for every user. ``` ```hint When things go wrong Consider how an edit, especially switching off a broken feature, reaches every running service, and what a service does when it cannot reach the flag system at all. ``` ### Clarifying Questions - Which callers check flags: backend services only, or also browsers and mobile apps? - Roughly how many flags, services and flag checks per second, and how much latency may one check add? - How quickly must a change, especially an emergency switch-off, reach every service? - Must a user's assignment stay the same across requests, services and devices, and what identifies a user? - Is this only gating and gradual rollout, or also multi-variant experiments that need exposure logging? - When the flag system is unreachable, should a check fail closed, fail open, or fall back to a per-flag default? - Are approvals, an audit history, or scheduled changes required for flag edits? ### What a Strong Answer Covers - Requirements and the read/write asymmetry established first, with the design following from them rather than from a template for another problem - A flag data model covering on/off, percentage rollout, allow and block lists and variants, with an explicit evaluation order - A deliberate choice of where evaluation happens (inside each service or in a central evaluation service), justified by latency and availability - Deterministic, sticky bucketing that keeps assignments stable while a rollout grows and independent across flags - A propagation path from the control plane to every service, with versioning, validation before publication, and a bounded delay for emergency switches - Behavior when the control plane, storage or network fails, including defaults and the last known good configuration - Audit and revert, permissions, exposure logging for experiments, and visibility into which configuration version each service runs ### Follow-up Questions - A feature is causing errors in production. How does switching its flag off reach every running instance within seconds? - Mobile and browser clients must check flags too. What changes, given that you cannot ship every rule and allow list to an untrusted device? - A rollout grows from 10% to 20% of users. How do you guarantee the original 10% stay in, and how do you keep two experiments from always choosing the same users? - An allow list grows to millions of user IDs. How does that affect your design?

Overview: System design question: design a feature flag service, similar in spirit to Amazon's Weblab, that lets teams turn features on or off, roll them out to a percentage of users, and force them on or off with allow and block lists. It probes where flag checks run, how users are assigned consistently, and how changes reach every service safely.

Read the full Stripe Software Engineer interview experience this question came from

|Home/System Design/Stripe
Stripe logo
Stripe
Aug 30, 2026
hardSoftware EngineerTechnical ScreenSystem Design
0
0

Design a feature flag service, similar in spirit to Amazon's Weblab, for a company that runs many backend services. Engineers create and edit flags in a management console. For each flag they can turn the feature on or off for everyone, roll it out to a percentage of users, and force it on or off for specific users with an allow list and a block list. Application services check flags while they handle requests.

Clarifying Questions Guidance

  • Which callers check flags: backend services only, or also browsers and mobile apps?
  • Roughly how many flags, services and flag checks per second, and how much latency may one check add?
  • How quickly must a change, especially an emergency switch-off, reach every service?
  • Must a user's assignment stay the same across requests, services and devices, and what identifies a user?
  • Is this only gating and gradual rollout, or also multi-variant experiments that need exposure logging?
  • When the flag system is unreachable, should a check fail closed, fail open, or fall back to a per-flag default?
  • Are approvals, an audit history, or scheduled changes required for flag edits?

What a Strong Answer Covers Guidance

  • Requirements and the read/write asymmetry established first, with the design following from them rather than from a template for another problem
  • A flag data model covering on/off, percentage rollout, allow and block lists and variants, with an explicit evaluation order
  • A deliberate choice of where evaluation happens (inside each service or in a central evaluation service), justified by latency and availability
  • Deterministic, sticky bucketing that keeps assignments stable while a rollout grows and independent across flags
  • A propagation path from the control plane to every service, with versioning, validation before publication, and a bounded delay for emergency switches
  • Behavior when the control plane, storage or network fails, including defaults and the last known good configuration
  • Audit and revert, permissions, exposure logging for experiments, and visibility into which configuration version each service runs

Follow-up Questions Guidance

  • A feature is causing errors in production. How does switching its flag off reach every running instance within seconds?
  • Mobile and browser clients must check flags too. What changes, given that you cannot ship every rule and allow list to an untrusted device?
  • A rollout grows from 10% to 20% of users. How do you guarantee the original 10% stay in, and how do you keep two experiments from always choosing the same users?
  • An allow list grows to millions of user IDs. How does that affect your design?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...