Debug a Google Cloud Web App Where 80% of Users See Only a Loading Spinner
Company: Costco
Role: Software Engineer
Category: Software Engineering Fundamentals
Difficulty: medium
Interview Round: Technical Screen
A product with a web frontend is deployed on Google Cloud. About 20% of users see the page load normally. The other 80% see only a loading spinner, and the page content never appears. Walk through how you would debug this.
The interviewer asked for a step-by-step walkthrough of the debugging process rather than an early narrowing to one suspected cause, such as a recent configuration or deployment change. Show how you would find where the failure is, and how you would confirm it.
```hint Read the spinner
The spinner itself is evidence. Work out what must already have loaded for it to appear, and what still has to happen before it goes away.
```
```hint What splits 80 from 20
A stable split of users rarely happens by chance. List the parts of a cloud deployment that can route or treat users differently, and decide how you would test each one.
```
### Constraints and Clarifications
- The setup behind the frontend is not specified: the APIs and backend services, data stores, CDN, frontend framework, and which Google Cloud services host it. Ask, or state your assumptions.
- Whether the problem started at a specific time, for example with a release, is not established. Treat that as something to find out during debugging, not as a given.
### Clarifying Questions
- Is the split stable per user (the same users always fail) or per page load (any user fails about 80% of the time)?
- Do the failing users share anything: browser, device, region, network, logged-in state, account type?
- What does the frontend load or request before it removes the spinner: which scripts, configuration and API calls?
- How is the app hosted and fronted: static hosting behind a CDN, App Engine, Cloud Run, GKE, a load balancer?
- What monitoring exists today: client-side error reporting, real-user monitoring, load balancer and backend logs, traces?
### What a Strong Answer Covers
- Scoping first: confirming the failure, and measuring who fails and whether the split is per user or per request
- Using what the spinner proves to narrow the failure to one stage of the page load
- Browser-side evidence from the network and console views of a failing load
- Server-side and infrastructure evidence across load balancing, backend instances or versions, caching and dependencies
- Candidate explanations for the 80/20 split, each tested and eliminated with specific evidence
- Limiting user impact during the investigation, fixing the root cause, and preventing a recurrence
### Follow-up Questions
- You cannot reproduce the problem on your own machine. How do you get evidence from affected users?
- Load balancer logs show that every failing request was served by the same subset of backend instances. How do you confirm the cause and fix it safely?
- Why might the page show a spinner forever instead of an error message, and what would you change in the frontend?
- How would you have detected this before users reported it?
Overview: A production debugging scenario: a web app deployed on Google Cloud loads for only 20% of users while the other 80% see a loading spinner forever. It tests a systematic walkthrough that scopes the failure, gathers browser and server evidence, and tests explanations for the split one at a time.
Read the full Costco Software Engineer interview experience this question came from