AE: The euro-chat question — there's a chatbot used for B2C customer service, and the question was how you'd define success, testing, etc.
Behavioral: Standard questions.
SQL: The shop visibility question. Overall pretty standard questions, and the last part asked me to define my own metric to prove whether new shops are more active. Afterward the recruiter's feedback said I had a red flag on this round — I think it's because my metric definition and approach were different from what the interviewer had in mind. My metric was probably a bit oversimplified, but I was already pressed for time answering everything within the limit, and I deliberately didn't design something too complex, otherwise I wouldn't have finished.
AE: Whether a search feature is successful is evaluated along 2 dimensions (relevancy, accuracy), both binary flags (1/0). The first few sub-questions were easy, but there was one later question I didn't answer well (I misremembered the formula, ugh) — it said there are two models handling this feature, each used by 100 people (1 use per person), one had 90 people report it as successful, the other had 85 people report it as successful (the exact numbers might be off, but that's the gist) — can you say one model is better than the other? Validate it using the known data.
Discussion
Loading comments…