Snowflake Software Engineer Interview Experience — Rate Limiter in CoderPad, Ran Out of Time Before the Distributed Follow-up

Snowflake·Software Engineer·Aug 2026
Technical ScreenRejectedhard

Snowflake phone screen, one hour, writing Python in CoderPad, and I was allowed to use AI. The recruiter said this round looks at how you ask about requirements, how you test, and how you debug.

[Question]

Build a rate limiting service. Different clients can have different rate limit tiers, for example free, pro and enterprise. The limits for each tier can be updated at runtime without restarting the service. A client makes a request with an identifier, and the service should enforce the appropriate limit for its tier. If a client exceeds its limit, the service should reject the request.

Verbal addition: if a client has no tier configured, or the config is broken, always fall back to a global default policy. Default deny or default approve are both fine.

[Follow-ups, in order]

  1. Why did you pick this rate limiting algorithm? What are the pros and cons of the different algorithms?
  2. After writing it: do you want to add some tests to verify it actually works?
  3. When testing, the real clock runs too fast and the bucket does not have time to refill. How do you test that? (Direction: mock the clock, or add a time.sleep between calls.)
  4. Closing question: right now it is a single instance with all state in memory. If this becomes a cluster of rate limiters, how would you evolve it into a distributed one?

[What it is testing]

It is not just about writing something that runs. The interviewer is watching how you choose an algorithm and explain the trade-offs, how you write tests, and how you handle time in tests. Changing limits at runtime is a trap: for example, if free goes from 5 to 10, the tokens in the bucket do not suddenly increase, capacity is only the upper bound. Also watch the refill formula: tokens to add = elapsed time * refill rate, not * capacity.

[Direction for the solution]

Token bucket algorithm: each client keeps a bucket that records the current token count and the last refill time. When a request comes in, first compute the elapsed time, add tokens at the refill rate (capped at capacity), then take 1 token. If there are not enough tokens, reject. Store the tier config in a dict, and changing the dict at runtime is enough, since the next request picks up the new capacity automatically. Clients without a configured tier go through the global fallback. For tests, use a mock clock or dependency injection to control time.

I felt that at the end the interviewer still wanted to ask about the distributed evolution, which is more of a system design direction, but I had no time left after finishing the coding, and I did not pass the interview in the end. Everyone should watch their time on this question.

Direction for the distributed evolution: store the bucket state in Redis and use a Lua script to guarantee atomicity.

Published

Curated and edited by PracHub

Practice the questions from this interview

Discussion

Sign in to join the discussion. The author is notified of every comment.

Loading comments…

Interview at a glance

Company
Snowflake
Role
Software Engineer
Rounds
Technical Screen
Outcome
Rejected
Difficulty
hard
Interview date
Aug 2026
Questions from this interview
1 question

Real Snowflake interview experiences

First-hand reports from Snowflake candidates — the rounds, the questions they were asked, and how it went.

All 21 Snowflake interview experiences