Anthropic Software Engineer Interview Experience — Messaging at 100 Million DAU and Gaps in My Design

Anthropic·Software Engineer·Sep 2026
Technical ScreenRejectedhard

The question was to design a simple text-messaging app, with a fairly clear scope:
Only conversations between two people, no group chats.
Text only, no images or videos.
Web clients only, with no mobile push notifications to consider.
100 million DAU.
Keep messages when the recipient is offline so they can be read after the recipient comes back online.
The interviewer asked me to share my entire screen, and I could choose my own drawing tool. The conversation was friendly overall. They gave hints and followed up on my design.
I started by listing requirements such as message persistence, ordering, deduplication, low latency, and scalability. The architecture was roughly a gateway, app servers maintaining WebSocket connections, a message store, and a presence service tracking whether users were online and which server they were on.
For online messages, persist first and then return success to the sender. Then find the recipient, forward the message to their app server, and push it through WebSocket. When offline users return, fetch missing messages using each conversation's last-seen sequence number.
It all sounded quite smooth. The problems only came out once the real follow-up questions started.
The follow-ups that stuck with me most:
With 100 million DAU, how exactly does your database handle the load?
I initially talked about adding app servers and read replicas. The interviewer brought the discussion back to the database: every message needs a write, and returning users need unread-message queries. Can those operations work at this scale?
After a reminder, I added a capacity estimate. With my own assumption of 100 messages sent per person per day, that's about 116,000 message writes per second on average, before accounting for peaks. Only then did we discuss sharding by chat ID.
My problem was that although I'd said at the beginning that I would do capacity estimation, I hadn't used numbers to support the storage and scaling plan first. Adding them after being pressed disrupted my flow.
When a user comes online, how do you know which chats they belong to?
I'd already explained using the chat ID, latest sequence number, and user's last-seen value to calculate unread messages. The interviewer then asked: how do you actually find all this user's chat IDs?
My answer was vague. I said we'd scan the conversations the user belonged to, but I didn't clearly explain the user-to-chat mapping, indexes, or specific queries.
Looking back, this was the most obvious gap: I knew what needed to be queried, but hadn't nailed down which key to use, which table to query, or how to retrieve the results.
Can you walk through a message's entire lifecycle?
The interviewer asked me to walk through the online and offline paths separately, asking along the way how the gateway chooses an app server, when success is returned to the sender, and how the recipient is located.
Then they asked what happens if a message is lost while being forwarded between two app servers, or if messages arrive out of order.
I answered that each chat has an increasing sequence number and the message store is the basis for recovery. If a client receives 10 and then jumps straight to 12, it knows 11 is missing and fetches it from storage. The relationship between persistence, real-time push, and fetching missing messages needs to be clear here.
If you use read replicas, how do you meet real-time requirements?
As soon as I proposed read replicas, the interviewer asked about the relationship between replication lag and real-time messages.
I distinguished the read scenarios: history and catching up after coming back online can tolerate some delay, while the latest missing messages that need immediate retrieval in an online conversation go to the primary.
My takeaway is that after proposing a component, I need to explain which requests use it and how it affects the original requirements.
If traffic suddenly surges, what breaks first? Why would adding a cache help?
I said we'd prioritize message persistence and online delivery, degrade historical-message loading, add app servers, and cache presence and message reads.
The interviewer then asked: how exactly does the cache help the message store?
Looking back, this part of my answer was still rather generic. I should first have clarified whether the pressure came from connection counts, writes, historical reads, or presence updates, then explained which load each measure could relieve.
Deduplication was another area I felt afterward that I hadn't fully explained. I said the client would send an idempotency token and the app server would use a small buffer for deduplication. But if the server restarts or a retry goes to another machine, that guarantee is incomplete. This is my own retrospective, not feedback from the interviewer.
There was also a very practical little issue in this round: I wasn't familiar enough with the drawing tool and spent time adjusting shapes. The interviewer even reminded me that the diagram didn't need to be so perfect. It's definitely worth practicing beforehand.
During my questions at the end, we talked about how engineers work. The interviewer mentioned that their own recent work involved more managing of agents, and they felt the team emphasized execution and had relatively aligned goals across teams. That was just one interviewer's personal impression at the time, for reference.

Published

Curated and edited by PracHub

Practice the questions from this interview

Discussion

Sign in to join the discussion. The author is notified of every comment.

Loading comments…

Interview at a glance

Company
Anthropic
Role
Software Engineer
Rounds
Technical Screen
Outcome
Rejected
Difficulty
hard
Interview date
Sep 2026
Questions from this interview
1 question

Real Anthropic interview experiences

First-hand reports from Anthropic candidates — the rounds, the questions they were asked, and how it went.

All 43 Anthropic interview experiences