As a Data Scientist at Atlassian, you play a central role in driving strategy, shaping product roadmaps, and unlocking the potential of millions of teams worldwide. Your work directly influences how flagship products like Jira, Confluence, and Trello evolve to serve modern distributed workforces. By turning massive volumes of telemetry, user behavior, and operational data into clear strategic insights, you empower engineering, product, and leadership teams to make confident, data-informed decisions.
This role sits at the intersection of rigorous analytics, product sense, and scalable engineering. You will tackle complex problem spaces ranging from growth loops and product adoption to organizational network analysis and Generative AI application performance. Whether you are designing sophisticated experimentation frameworks or diagnosing unexpected metric shifts in core user funnels, your contributions directly impact business growth and user experience.
The environment at Atlassian is fast-paced, collaborative, and deeply analytical, requiring you to thrive in ambiguity and structure open-ended challenges into actionable roadmaps. You will operate as a trusted partner to cross-functional stakeholders, translating complex quantitative findings into compelling narratives that mobilize action. Expect a culture that values intellectual curiosity, rigorous methodology, and a relentless focus on unleashing the potential of teams.
Recruiter Conversation
reportedA screening call is a matching exercise run by someone who will not evaluate your statistics. They are checking that the work described on your resume is work you personally did, and that its scope matches the level the role is written for. Logistics get settled in the same half hour so nobody spends an interviewer's afternoon on a mismatch. The answer that fails is the one narrated in the plural. If every sentence is 'we built' and 'the team decided', there is nothing specific to write down about you. Name the piece that was yours, the decision you made inside it, and what changed after.
What to demonstrate
- Whether the ownership implied by your resume survives one round of follow-up about who actually did which part
- Whether your described scope (data size, stakeholders, what shipped) matches the seniority the role is written at
- Whether timeline, location and compensation expectations make the rest of the loop worth scheduling
How to prepare
- Rewrite your top three resume bullets in the first person singular, each with the decision you made and what moved afterwards, then say them out loud once so the 'we' does not return under pressure
- Attach one number to each project: the baseline, the change, and the window it was measured over. Where impact was never measured, say that plainly rather than inventing a figure
- Settle your compensation range before the call and give it as a range with a reason behind it, such as current total comp or a competing timeline, instead of deflecting the question twice
Online Assessment
reportedBefore anything else, this round is a reading test. You are given a small schema and a question phrased in business language, and most of the difficulty sits in the gap between them. Who counts as an active user, does a refunded order still count as an order, is that date column an event time or a load time. Weak answers start typing immediately and compute something precise about the wrong population. Strong ones pin the definition in one sentence, name the column that encodes it, then write the query. On a timed assessment with nobody to tell, write the definition in a comment anyway.
What to demonstrate
- Whether an ambiguous term becomes a specific column and filter before any computation happens
- Whether you read the schema for keys and cardinality rather than only for column names
- Whether the result answers the question at the grain it was asked at, per user or per session or per day
How to prepare
- Take three metrics you already use and write down the exact filter and exact grain behind each, then practise stating one of them in a single sentence out loud
- On a schema you have never seen, spend the first minute writing what one row of each table means and which key it is unique on, then predict which joins can duplicate rows
- Rehearse a version where the definition changes halfway through, and edit the query you have instead of starting over
Onsite Loop
reportedWhere a loop includes a partner from outside the data team, that conversation usually carries the same weight as the technical ones and gets the least preparation. The person opposite you will not follow a derivation and does not need to. They are working out whether having you involved would make their decisions better or slower. The failure mode is not being too technical. It is answering a question about a decision with a description of your method, leaving the translation to them. What they carry into the debrief is the sentence you handed them, not the analysis underneath it.
What to demonstrate
- Whether a statistical result arrives as something the partner could act on, with the one caveat that would change their decision kept and the rest left out
- Whether you can state what you need from their side, in their terms: instrumentation that does not exist yet, a definition they own, or a holdout they have to agree to
- Whether uncertainty is given as a range someone can plan against, rather than as hedging that invites them to ignore the result
- Whether you ask what decision is actually on the table before explaining anything
How to prepare
- Take a result you know well and write the version for someone who stops reading after one sentence, then the three-minute version, and check the short one is not the long one with the qualifications stripped out
- For a past project, list everything you asked a non-technical partner for and how you phrased it, then rewrite each ask so it names what goes unmeasured without it
- Practise saying where a result does not apply, out loud, in one sentence that a partner could repeat accurately to someone else
PracHub editorial advice for the preparation topics above.
Comparing accounts that received a sales or customer-success touch against those that did not
Assignment of coverage is deliberate and pulls in both directions at once: the largest accounts get a named owner because they are valuable, and the accounts showing distress get one because they are at risk. The comparison therefore mixes a strong positive selection with a strong negative one, and the naive estimate can come out with either sign depending on which assignment rule dominated during the period examined. Nothing about matching on observed size fixes this, because the risk signal that triggered coverage is usually the same signal that predicts the outcome. It needs either an actual randomised or staggered rollout of coverage, or a design built on a capacity constraint or territory boundary that assigns coverage for reasons unrelated to account health.
Treating raw request or usage volume as engagement
Most traffic in this domain is emitted by machines. Continuous-integration pipelines, scheduled batch jobs, synthetic monitors, backfills and client retries can all grow by an order of magnitude from one configuration change made by one engineer, and none of it represents a new decision to use the product. The inversion is what makes it dangerous: when the platform degrades, clients retry, so error-driven retry volume rises at the exact moment the customer is most likely to leave, and an engagement dashboard built on raw counts shows growth immediately before a churn. Filter on traffic_class and on successful status before anything else, and keep failed-request volume as its own separate series.
Reading experiment results before checking the arm split
Compare observed arm counts against the intended allocation ratio, not an assumed even split, and set the alarm far below the conventional 0.05: at 0.05 roughly one healthy experiment in twenty trips it, which is why sample-ratio checks usually run at p < 0.001 or stricter. The test's power scales with sample size, so it misses a real diversion on a small experiment and fires on an imbalance too small to move the estimate on a very large one. A flag means go find the assignment or logging fault before reading any outcome, not report a mismatch.
Never asking what decision the analysis will inform
Open with who makes the decision, what the options are, and by when. The answer determines the precision you need, the segments worth cutting, and whether an observational read suffices or an experiment is required.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
How do you test for normality, and what non-parametric alternatives do…
How do you test for normality, and what non-parametric alternatives do you use when your data violates standard distribution assumptions?
Approach
- Write down the assumption the method needs before you use the method.
- Quantify uncertainty explicitly rather than reporting a point estimate alone.
- Say what the estimate is of, and over what population it generalises.
Follow-up
- Which assumption here is most likely to be violated in practice?
- How would you explain this result to someone who does not know statistics?
How do you monitor model drift and data degradation after deploying a …
How do you monitor model drift and data degradation after deploying a machine learning model to production?
Approach
- Say how the offline result would be validated online before it is trusted.
- Pick an evaluation metric that matches the cost of each error type, not a default.
- Set a baseline first, so any model has something honest to beat.
Follow-up
- How would you choose the decision threshold, and who owns that choice?
- Where could label leakage enter this setup?
Sessionise an API event stream with a 30-minute inactivity gap
fct_api_request arrives as a DataFrame with account_id, user_id, request_at (tz-aware UTC), traffic_class and http_status, roughly 5 million rows. Assign a session_id to every human-attributable request: drop rows where user_id is null or traffic_class is in ('ci','synthetic_monitor','load_test'), then open a new session whenever the gap since that user's previous remaining request exceeds 30 minutes. Return the filtered frame plus session_id, and a per-session summary with user_id, account_id, session start, session end and request count. Do not loop over rows.
Approach
- Settle the filter-then-gap ordering before writing code. Removing CI and synthetic rows changes the gaps, so sessionising the raw stream and filtering afterwards is a different answer; the definition given filters first, and the two diverge most for accounts whose CI runs every ten minutes.
- Sort once by (user_id, request_at) with a stable kind, then gap = df.groupby('user_id', sort=False).request_at.diff(). The first row of each user yields NaT, which is exactly the boundary condition you want rather than a special case to patch.
- new_session = gap.isna() | (gap > Timedelta(minutes=30)); session_id = new_session.cumsum(). The cumsum runs over the whole sorted frame and therefore produces globally unique ids in one pass; a per-user cumcount collides across users and forces a composite key on every downstream join.
- Build the summary with a single groupby('session_id').agg(...). user_id and account_id can be carried with 'first' only because the sort key groups them — state that dependency, since it silently breaks if someone later re-sorts the frame.
- Decide explicitly what a session means when one user_id holds memberships in several accounts: either add account_id to the sort and group keys, or document that sessions may cross accounts. Leaving it undecided produces sessions whose account_id is whichever row sorted first.
Worked solution 30 min
- human = df[df.user_id.notna() & ~df.traffic_class.isin(['ci','synthetic_monitor','load_test'])].sort_values(['user_id','request_at'], kind='mergesort').reset_index(drop=True)
- gap = human.groupby('user_id', sort=False).request_at.diff(); human['session_id'] = (gap.isna() | (gap > pd.Timedelta(minutes=30))).cumsum()
- summary = human.groupby('session_id').agg(user_id=('user_id','first'), account_id=('account_id','first'), start=('request_at','min'), end=('request_at','max'), n_requests=('request_at','size')).reset_index()
- assert summary.n_requests.sum() == len(human) and human.groupby('session_id').user_id.nunique().max() == 1
Follow-up
- Where does 30 minutes come from, and how would you pick it from this data instead of from convention?
- An engineer reused their personal key for a nightly batch job, so machine traffic carries a human user_id. How would you detect that, and should those requests form sessions?
- How much does the session count change if you sessionise before dropping CI traffic rather than after?
Write a SQL query using window functions to calculate the rolling 30-d…
Write a SQL query using window functions to calculate the rolling 30-day active user count for each workspace.
Approach
- Say which table is the grain you start from, and join outward from it.
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- Check whether any join is one-to-many before aggregating, or the sums inflate.
Follow-up
- What breaks if events arrive late or out of order?
- How does the query change if the join becomes one-to-many?
Given a user event log table, write a query to identify multi-step con…
Given a user event log table, write a query to identify multi-step conversion drop-offs across the onboarding funnel.
Approach
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
- Say which table is the grain you start from, and join outward from it.
- State the window function and its partition and ordering out loud before writing it.
Follow-up
- How does the query change if the join becomes one-to-many?
- What breaks if events arrive late or out of order?
Sessionise interactive API traffic with a thirty-minute inactivity gap
fct_api_request carries account_id, user_id (null for service accounts), request_at, traffic_class and http_status. Using only traffic_class = 'interactive' rows with a non-null user_id, group each seat's requests into sessions on a 30-minute inactivity threshold: a request more than 30 minutes after the previous request from the same (account_id, user_id) opens a new session. For one ISO week return, per account, the session count, the median session duration in minutes and the median requests per session. A single-request session has a duration of zero.
Approach
- Filter first: traffic_class = 'interactive' and user_id IS NOT NULL. Machine traffic has no sessions in any useful sense, and leaving CI or batch rows in produces sessions that are really cron schedules.
- Get the previous timestamp with LAG(request_at) OVER (PARTITION BY account_id, user_id ORDER BY request_at). Partitioning by user_id alone stitches one person's work across two different accounts into one fabricated session, because a human holds memberships in several accounts.
- Flag a boundary where the lag is NULL or request_at - lag > interval '30 minutes'. Decide and state whether exactly 30 minutes continues the session; strictly greater is the conventional choice and needs to be written down either way.
- Assign session ids with SUM(boundary::int) OVER (PARTITION BY account_id, user_id ORDER BY request_at ROWS UNBOUNDED PRECEDING), the standard running-count construction for islands.
- Roll up to sessions with min(request_at), max(request_at) and count(*), then to accounts with percentile_cont(0.5) WITHIN GROUP (ORDER BY ...). Use medians, not means: session length is strongly right-skewed and one long-running client dominates the average.
Worked solution 30 min
- Build the filtered week: interactive rows with user_id IS NOT NULL inside [week_start, week_start + 7 days).
- Add prev_at via LAG(request_at) OVER (PARTITION BY account_id, user_id ORDER BY request_at) and derive is_new_session = (prev_at IS NULL OR request_at - prev_at > interval '30 minutes').
- Add session_seq = SUM(is_new_session::int) OVER (PARTITION BY account_id, user_id ORDER BY request_at ROWS UNBOUNDED PRECEDING).
- Aggregate to sessions by (account_id, user_id, session_seq) taking min, max and count, with duration_minutes = EXTRACT(EPOCH FROM max - min) / 60.
- Aggregate to accounts with count(*) AS sessions, percentile_cont(0.5) WITHIN GROUP (ORDER BY duration_minutes) and percentile_cont(0.5) WITHIN GROUP (ORDER BY request_count).
Follow-up
- Sessions belonging to colleagues in one account are correlated. What does that do to a t-test on session length across an experiment arm?
- A long-poll or streaming endpoint keeps a connection open for hours. How do you stop it reading as one twelve-hour session?
A core engagement metric dropped by 15% week-over-week. How would you …
A core engagement metric dropped by 15% week-over-week. How would you investigate and diagnose the root cause?
Approach
- Name one primary metric, then the guardrail that stops it being gamed.
- State what result would change your recommendation, so the answer is falsifiable.
- Decompose the metric into the rates that drive it, and say which one you would check first.
Follow-up
- Which segment would you cut first, and what would that rule out?
- What would you do if the primary metric and the guardrail moved in opposite directions?
How would you construct a composite health score for enterprise accoun…
How would you construct a composite health score for enterprise accounts using product usage and support ticket data?
Approach
- Restate the decision this analysis has to support, and who acts on the answer.
- State what result would change your recommendation, so the answer is falsifiable.
- Name one primary metric, then the guardrail that stops it being gamed.
Follow-up
- How would you detect that the metric is being gamed rather than genuinely improving?
- What would you do if the primary metric and the guardrail moved in opposite directions?
What statistical methods would you use to measure user retention when …
What statistical methods would you use to measure user retention when observation windows vary significantly across cohorts?
Approach
- Restate the decision this analysis has to support, and who acts on the answer.
- Decompose the metric into the rates that drive it, and say which one you would check first.
- State what result would change your recommendation, so the answer is falsifiable.
Follow-up
- What would you do if the primary metric and the guardrail moved in opposite directions?
- Which segment would you cut first, and what would that rule out?
Share an example of how you prioritized multiple competing requests fr…
Share an example of how you prioritized multiple competing requests from different engineering and product leads.
Approach
- State what result would change your recommendation, so the answer is falsifiable.
- Name one primary metric, then the guardrail that stops it being gamed.
- Fix the population and the time window before naming any metric.
Follow-up
- Which segment would you cut first, and what would that rule out?
- How would you detect that the metric is being gamed rather than genuinely improving?
Explain how you handle sample ratio mismatch (SRM) and what steps you …
Explain how you handle sample ratio mismatch (SRM) and what steps you take to debug the underlying data pipeline issue.
Approach
- Decide the analysis before seeing data, including how long it runs and when you look.
- Name the guardrails that would stop a launch even on a positive primary result.
- Say whether units interfere with each other, and switch design if they do.
Follow-up
- How would you handle interference between treated and control units?
- What would you conclude if the result is positive but the test is underpowered?
When would you choose a Bayesian approach over a frequentist approach …
When would you choose a Bayesian approach over a frequentist approach for evaluating an experiment?
Approach
- State the primary metric and the minimum effect worth shipping, then size the test.
- Decide the analysis before seeing data, including how long it runs and when you look.
- Name the guardrails that would stop a launch even on a positive primary result.
Follow-up
- How would you handle interference between treated and control units?
- What would you do if you could not randomise at all?
How would you evaluate whether a new freemium tier upgrade path is can…
How would you evaluate whether a new freemium tier upgrade path is cannibalizing enterprise conversions?
Approach
- Work from the decision backwards to the evidence you would need.
- State your assumptions explicitly before working the problem.
- Say what you would check first and why it is the highest-information step.
Follow-up
- How would you know your answer was wrong?
- What assumption would you test first?
Measure switching cost when switching cost is not observable
Leadership wants switching cost tracked as a leading indicator of renewal. Nothing in the warehouse records it. Available: fct_api_request (account_id, workspace_id, environment, api_key_id, sdk_name, sdk_version, endpoint, traffic_class, http_status, request_at) and dim_account (account_id, employee_band, deployment_model, is_internal). Propose a proxy, state the direction and likely size of its bias, and name one decision the proxy is good enough for and one it is not. Deliver the proxy definition, the written bias statement, and the validation you would run against observed renewal outcomes.
Approach
- Say first that switching cost is unobservable in this data and that the deliverable is a biased correlate with its bias written down, not a measurement. Anything presented as a direct measure of switching cost is already wrong before the SQL starts.
- Define the proxy as production integration breadth per account over a trailing 28 days: distinct endpoint route templates, distinct api_key_id, and distinct workspace_id with environment = 'production', all restricted to http_status < 400 and traffic_class in ('interactive','batch'). Breadth, not volume, because request volume is one CI configuration change away from an order of magnitude.
- State the biases with their mechanisms and their sign. Upward with account size, because breadth correlates with employee_band, so the proxy ranks large accounts as sticky whether or not they depend on anything. Downward for deployment_model = 'self_hosted', whose traffic does not all cross the managed gateway, so their breadth is systematically understated. Blind to criticality: one endpoint carrying a production billing path is a larger switching cost than twenty endpoints behind a read-only dashboard, and nothing in this data distinguishes them.
- Handle the bias by stratifying rather than by pretending it is gone. Report the proxy within employee_band and deployment_model strata and state explicitly that cross-stratum comparisons are not supported by the construction.
- Validate against the only outcome that is actually observed: renewal on the renewal-eligible base. Within strata, report renewal rate by proxy quintile for accounts whose term_end_date has passed plus a 45-day grace, and report discrimination at the operating point a capacity-bound team can work rather than a global AUC.
Worked solution 30 min
- Compute the trailing-28-day proxy per account: distinct endpoint, distinct api_key_id and distinct production workspace_id under successful, interactive-or-batch filters.
- Join dim_account for employee_band and deployment_model and drop is_internal = true.
- Build the renewal-eligible cohort from fct_subscription_period term_end_date in the target months with a 45-day grace, labelling each account renewed or not.
- Within each employee_band stratum, report renewal rate by proxy quintile, and report the proxy distribution for self_hosted accounts separately so its understatement is visible.
Follow-up
- An account's endpoint breadth drops 40% in a week. Name three explanations that have nothing to do with reduced dependency.
- Which decision would you refuse to make on this proxy, and what evidence would you need before making it?
Billable units per account jumped while nothing shipped
Billable units per paying account rose 22 percent month over month with no pricing or packaging change. From fct_api_request (account_id, endpoint, http_status, is_retry, idempotency_key, traffic_class, billable_units, request_at) and fct_usage_daily (account_id, sku_code, usage_date, billable_quantity), determine how much of the rise is delivered value and how much is duplicated work. Deliverable: the decomposed figure, the accounts it concentrates in, and a recommendation on whether to report the 22 percent at all.
Approach
- Check the guardrail before the headline: compute the customer-visible server error rate per account for both months, numerator http_status >= 500 and denominator excluding synthetic_monitor and load_test. Metered volume rising alongside an error rate is the known failure mode in this domain.
- Deduplicate logical work by counting billable_units once per (account_id, idempotency_key) at the first successful request rather than once per row. Rows with a null idempotency_key cannot be deduplicated, so report their share as an explicit uncertainty band instead of assuming they are all unique.
- Split by traffic_class before interpreting anything, because one change to a continuous-integration configuration can multiply request volume overnight without a human deciding anything about the product.
- Test concentration: compute the per-account distribution of the increase and its top-decile share. A rise carried by a few accounts scaling one batch job is a different finding from a broad shift and gets a different recommendation.
- Reconcile against fct_usage_daily for the same accounts and dates, and be ready to explain the expected gap in two sentences: request rows include retries and failures carrying zero billable_units, and the usage table restates after first write.
Follow-up
- A client retrying a request the server already completed produces duplicate billed work. What protocol or product change removes that, and what would you measure to confirm it worked?
- If the duplicated volume was genuinely invoiced, what should finance do, and how does that change what belongs in the metric?
Roughly 90 minutes a night on weekdays with one longer weekend block. The plan deliberately cuts scope rather than compressing everything, on the assumption that finishing one thing a night beats half-starting four.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Fix the scope and set a baseline
- Read the role description and write the three things the loop will almost certainly test, then write an explicit not-doing list for everything else and keep it visible all week.
- Take one 20-minute SQL prompt and one 10-minute metric question cold, and write the single sentence that says what blocked each attempt, since that sentence is what decides which two topics get the most evenings.
- Set the week's one rule: one problem finished to completion every night, including the night you only have 40 minutes.
Deliverable: A one-page scope with an explicit not-doing list and two cold attempts, each carrying one sentence on what blocked it.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02One query pattern, written three times
- Choose the single pattern most likely to appear (a cohort retention grid, or a funnel counted by user) and write it three times from a blank file rather than editing the previous attempt.
- On the third attempt, write the grain of every CTE as a comment before writing its body.
- Stop at 90 minutes even if the third version is imperfect, and write the one thing you would fix with another hour.
Deliverable: Three independent versions of the same query plus a note on what changed between them.
Practice prompt ↗Practice prompt ↗Practice prompt ↗03Only the statistics you will be asked to defend
- Write, in under 200 words, how you would decide whether a difference between two groups is real: the test, its assumptions, and what you would switch to when an assumption fails.
- Compute a 95 percent confidence interval for a difference in proportions by hand on realistic numbers, then write in one sentence what changes if the two samples are paired rather than independent.
- Write your answer to "what does a p-value mean", check it against a definition, and delete the version that describes it as the probability the hypothesis is true.
Deliverable: A 200-word written answer and one hand-computed interval you can reproduce under pressure.
Practice prompt ↗Practice prompt ↗Practice prompt ↗04One case, and the assumptions holding it up
- Answer one product case aloud in 20 minutes with a recording running, then listen back with a pen and mark every claim you asserted without saying what it rested on: an assumed user behaviour, an assumed data source, an assumed baseline rate, an assumed grain.
- Pick the three assumptions the recommendation actually depends on, write how you would check each one against data, and say which one being wrong would flip the recommendation rather than merely weaken it.
- Write the four-step structure you used onto a card small enough to hold in working memory when you are nervous.
Deliverable: One recording, three load-bearing assumptions each with a written check, and a four-step structure card.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗05Your own work, timed
- Write a 90-second version and a four-minute version of your main project, and time both out loud rather than reading them.
- Prepare answers to the two follow-ups that always come: what you would do differently, and how you knew it worked.
- Put one number in the first sentence and be able to say exactly where that number came from and what it excludes.
Deliverable: Two timed narratives with one defensible number in the opening line.
Practice prompt ↗Practice prompt ↗06The one full rehearsal, in a longer weekend block
- Run a 60-minute mock covering query work, a case and a behavioural question in a single sitting with no breaks, because sustained attention is the thing evenings have not trained.
- Immediately afterwards, and before hearing any feedback, write the three moments you lost the thread.
- Spend the rest of the block only on those three moments, and on nothing you merely feel shaky about.
Deliverable: Mock notes naming three failure moments with a specific fix written under each.
Practice prompt ↗Practice prompt ↗07Taper
- Write the 20-minute warm-up you will actually do on the morning of the interview: one query you can already write from a blank file, one metric you can define out loud, and nothing you have never seen before.
- Re-read only your own notes from this week, and open no new material.
- Write down the logistics: the tool you will be asked to work in, whether lookups are allowed, and the sentence you will use when you do not know something.
Deliverable: A one-page card holding the case structure, the project numbers, and the logistics.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
An answer without a quantity is hard to interrogate, so interviewers keep probing until they find one. Come with the baseline, the change, the window it was measured over, and how confident you were. If the effect never got measured, say so and say what you would have measured. Fabricated precision is worse than an honest gap.
Tell me about a time when you had to influence a cross-functional prod…
Tell me about a time when you had to influence a cross-functional product roadmap using data, despite pushback from stakeholders.
Approach
- State the situation in two sentences and spend the rest on your reasoning.
- Close with what you would do differently, concretely.
- Pick a story where you drove the decision, not one where you observed it.
Follow-up
- What would you do differently if you ran that project again?
- How did you know the outcome was caused by your change?
Announce a metric fix that cuts the headline number
Weekly active organisations, the count on the company dashboard, has never excluded rows where dim_account.is_internal is true, and it counts traffic with traffic_class in synthetic_monitor and load_test. Correcting both reduces that count by 11 percent and removes most of the growth reported over two quarters. The figure appears in a board deck and in two teams' quarterly goals, one written on the count and one on the weekly active organisation ratio, whose denominator is accounts whose account_status was in ('trial','free','active_paid') through the week. Decide the order in which you tell people, what the dashboard shows during the transition, and what you propose happens to goals already set against the old definition.
Approach
- The interviewer is probing whether you can land a correction as an operational change with a plan attached, rather than as an announcement other people then have to clean up after.
- Quantify each exclusion separately before telling anyone: internal accounts, synthetic monitors, load tests. Three known quantities are a discussion; one alarming total is an argument.
- Be precise about which side of the metric each exclusion touches, because one team's goal is on a count and the other's is on a ratio. The traffic-class filters remove requests, so they shrink the numerator only. Dropping internal accounts removes them from the ratio's denominator as well, since internal accounts carry ordinary account_status values and therefore sit in that denominator. Internal accounts are active in almost every week while the real base is not, so the numerator loses a larger share than the denominator and the ratio falls by less than the count does. Compute both and say which one the 11 percent is before anybody assumes.
- Check whether the trend changes, not only the level. A constant 11 percent shift is a rebasing and nothing more. A shift that widens over time means the reported growth was partly internal or synthetic, which makes the existing goals unachievable as written and changes what you are asking teams to do.
- Sequence the disclosure: the metric owner and the two teams whose goals move first and privately, then the board channel with a written bridge, then the dashboard. The dashboard is last because a number that changes without explanation is read as instability rather than as a fix.
- Run both series for one reporting period with the bridge visible, restate history rather than letting the series break at a date, and set the date the old series is removed.
- Propose the goal treatment yourself: rebase each target by the shift measured on the metric that target is written against, rather than leaving each team to negotiate individually, which is where corrections of this kind usually die.
Follow-up
- One team's quarterly goal is now unreachable. Rebase the target or let it miss, and what does each choice teach the organisation?
- How would this have been caught when the metric was first defined?
- What else on that dashboard shares this failure mode, and how would you find out this week?
Report an underpowered consumption test to a non-technical executive
An account-randomised packaging change ran six weeks across 900 paying accounts. The effect on billable units per account per month is plus 4.1 percent, with a 95 percent interval from minus 3.2 to plus 11.8 after clustering standard errors at the account and applying the pre-registered winsorisation at the 99th percentile. An executive with no statistical background wants one number this week to decide a full rollout. Produce a three-sentence spoken answer, one chart, and an explicit recommendation of ship, stop or keep running, with the cost of each option stated.
Approach
- The interviewer is probing whether you can be decision-useful without either hiding the uncertainty or hiding behind it. Start from the decision rather than the statistics: establish what the executive would do differently at plus 4 percent versus zero, because if the action is identical the interval does not matter.
- Translate the interval into consequences in units the executive already reasons about. Multiply both endpoints by the cohort's baseline consumption and contracted rates to give an annualised revenue range, so the answer is a range of dollars rather than a range of percentages.
- Price the option to wait. Using the observed variance, state roughly how many additional account-weeks halve the interval width, so keep running becomes a quantified choice instead of a stall.
- Offer a cheaper path to the same decision: a lower-variance proximate outcome such as successful billable units on the new SKU, or CUPED using each account's pre-period consumption, quoting the expected variance reduction as one minus the squared pre-post correlation.
- Give a recommendation and name the single observation that would reverse it. A strong answer commits; a generic one recites the interval and leaves the decision on the table.
Follow-up
- The executive says it clearly works and is just not provable, so ship it. What is your answer?
- How much of the interval width comes from clustering and how much from the revenue tail, and what would you do about each?
- If you had to ship this week with no more data, which guardrail would you watch for the first fortnight and at what threshold would you roll back?
- 01
Tell me about a time when you had to influence a cross-functional product roadmap using data, despite pushback from stakeholders.
- 02
Weekly active organisations, the count on the company dashboard, has never excluded rows where dim_account.is_internal is true, and it counts traffic with traffic_class in synthetic_monitor and load_test. Correcting both reduces that count by 11 percent and removes most of the growth reported over two quarters. The figure appears in a board deck and in two teams' quarterly goals, one written on the count and one on the weekly active organisation ratio, whose denominator is accounts whose account_status was in ('trial','free','active_paid') through the week. Decide the order in which you tell people, what the dashboard shows during the transition, and what you propose happens to goals already set against the old definition.
- 03
An account-randomised packaging change ran six weeks across 900 paying accounts. The effect on billable units per account per month is plus 4.1 percent, with a 95 percent interval from minus 3.2 to plus 11.8 after clustering standard errors at the account and applying the pre-registered winsorisation at the 99th percentile. An executive with no statistical background wants one number this week to decide a full rollout. Produce a three-sentence spoken answer, one chart, and an explicit recommendation of ship, stop or keep running, with the cost of each option stated.
Is this an official Atlassian interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at Atlassian. Rounds and questions reflect what candidates have reported, not a process Atlassian has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How difficult is the interview process, and how much preparation time should I expect?
The interview loop is moderately to highly rigorous, emphasizing both technical depth and practical product judgment. Most candidates spend between four to six weeks in intensive preparation, focusing heavily on advanced SQL, experimentation design, and structured case studies.
PracHub interview research ↗What differentiates successful candidates from those who do not pass?
Successful candidates stand out by structuring ambiguous problems methodically before diving into math or code. They combine technical fluency with strong business acumen, demonstrating how their analytical recommendations directly drive product strategy and user value.
PracHub interview research ↗How should I approach the behavioral and values interview rounds?
Interviewers look for concrete examples of ownership, collaboration, and resilience, especially when navigating disagreement or project ambiguity. Ground your answers in past professional experiences, clearly highlighting your specific contributions and the ultimate business impact of your work.
PracHub interview research ↗Are interviews conducted remotely, and how does the distributed model affect the process?
All interview rounds are conducted virtually via video conferencing and shared online coding environments. Because the company operates as a distributed-first organization, demonstrating strong asynchronous communication skills and comfort with virtual collaboration is a distinct advantage.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22