Atlassian · Data Scientist
Updated · 2026-09-22

Atlassian Data Scientist
Interview Questions & Guide 2026

THE 60-SECOND BRIEF

As a Data Scientist at Atlassian, you play a central role in driving strategy, shaping product roadmaps, and unlocking the potential of millions of teams worldwide. Your work directly influences how flagship products like Jira, Confluence, and Trello evolve to serve modern distributed workforces. By turning massive volumes of telemetry, user behavior, and operational data into clear strategic insights, you empower engineering, product, and leadership teams to make confident, data-informed decisions.

Seniority shifts the scope more than the words in the title do. Earlier-career loops mostly check that you execute a well-posed analysis correctly; senior loops check that you can decide which question is worth answering and defend what you chose not to do.

Atlassian candidates report 3 rounds · ≈ 3-5 weeks. The stages below are what candidates describe, not a published process.

Read NRR on a fixed account cohortStrip CI, retry and synthetic traffic firstSeparate contracted seats from actively used seats

33 min read

Practice 18 Data Scientist prompts
12Company bank questionsSnapshot · Oct 2, 2026 PT
18Practice promptsAcross five skill areas
3With worked solutionsIncluded in the practice prompts

As a Data Scientist at Atlassian, you play a central role in driving strategy, shaping product roadmaps, and unlocking the potential of millions of teams worldwide. Your work directly influences how flagship products like Jira, Confluence, and Trello evolve to serve modern distributed workforces. By turning massive volumes of telemetry, user behavior, and operational data into clear strategic insights, you empower engineering, product, and leadership teams to make confident, data-informed decisions.

This role sits at the intersection of rigorous analytics, product sense, and scalable engineering. You will tackle complex problem spaces ranging from growth loops and product adoption to organizational network analysis and Generative AI application performance. Whether you are designing sophisticated experimentation frameworks or diagnosing unexpected metric shifts in core user funnels, your contributions directly impact business growth and user experience.

The environment at Atlassian is fast-paced, collaborative, and deeply analytical, requiring you to thrive in ambiguity and structure open-ended challenges into actionable roadmaps. You will operate as a trusted partner to cross-functional stakeholders, translating complex quantitative findings into compelling narratives that mobilize action. Expect a culture that values intellectual curiosity, rigorous methodology, and a relentless focus on unleashing the potential of teams.

01

Recruiter Conversation

reported

A screening call is a matching exercise run by someone who will not evaluate your statistics. They are checking that the work described on your resume is work you personally did, and that its scope matches the level the role is written for. Logistics get settled in the same half hour so nobody spends an interviewer's afternoon on a mismatch. The answer that fails is the one narrated in the plural. If every sentence is 'we built' and 'the team decided', there is nothing specific to write down about you. Name the piece that was yours, the decision you made inside it, and what changed after.

What to demonstrate

  • Whether the ownership implied by your resume survives one round of follow-up about who actually did which part
  • Whether your described scope (data size, stakeholders, what shipped) matches the seniority the role is written at
  • Whether timeline, location and compensation expectations make the rest of the loop worth scheduling

How to prepare

  • Rewrite your top three resume bullets in the first person singular, each with the decision you made and what moved afterwards, then say them out loud once so the 'we' does not return under pressure
  • Attach one number to each project: the baseline, the change, and the window it was measured over. Where impact was never measured, say that plainly rather than inventing a figure
  • Settle your compensation range before the call and give it as a range with a reason behind it, such as current total comp or a competing timeline, instead of deflecting the question twice
PracHub interview research ↗
02

Online Assessment

reported

Before anything else, this round is a reading test. You are given a small schema and a question phrased in business language, and most of the difficulty sits in the gap between them. Who counts as an active user, does a refunded order still count as an order, is that date column an event time or a load time. Weak answers start typing immediately and compute something precise about the wrong population. Strong ones pin the definition in one sentence, name the column that encodes it, then write the query. On a timed assessment with nobody to tell, write the definition in a comment anyway.

What to demonstrate

  • Whether an ambiguous term becomes a specific column and filter before any computation happens
  • Whether you read the schema for keys and cardinality rather than only for column names
  • Whether the result answers the question at the grain it was asked at, per user or per session or per day

How to prepare

  • Take three metrics you already use and write down the exact filter and exact grain behind each, then practise stating one of them in a single sentence out loud
  • On a schema you have never seen, spend the first minute writing what one row of each table means and which key it is unique on, then predict which joins can duplicate rows
  • Rehearse a version where the definition changes halfway through, and edit the query you have instead of starting over
PracHub interview research ↗
03

Onsite Loop

reported

Where a loop includes a partner from outside the data team, that conversation usually carries the same weight as the technical ones and gets the least preparation. The person opposite you will not follow a derivation and does not need to. They are working out whether having you involved would make their decisions better or slower. The failure mode is not being too technical. It is answering a question about a decision with a description of your method, leaving the translation to them. What they carry into the debrief is the sentence you handed them, not the analysis underneath it.

What to demonstrate

  • Whether a statistical result arrives as something the partner could act on, with the one caveat that would change their decision kept and the rest left out
  • Whether you can state what you need from their side, in their terms: instrumentation that does not exist yet, a definition they own, or a holdout they have to agree to
  • Whether uncertainty is given as a range someone can plan against, rather than as hedging that invites them to ignore the result
  • Whether you ask what decision is actually on the table before explaining anything

How to prepare

  • Take a result you know well and write the version for someone who stops reading after one sentence, then the three-minute version, and check the short one is not the long one with the qualifications stripped out
  • For a past project, list everything you asked a non-technical partner for and how you phrased it, then rewrite each ask so it names what goes unmeasured without it
  • Practise saying where a result does not apply, out loud, in one sentence that a partner could repeat accurately to someone else
PracHub interview research ↗

PracHub editorial advice for the preparation topics above.

01

Comparing accounts that received a sales or customer-success touch against those that did not

Assignment of coverage is deliberate and pulls in both directions at once: the largest accounts get a named owner because they are valuable, and the accounts showing distress get one because they are at risk. The comparison therefore mixes a strong positive selection with a strong negative one, and the naive estimate can come out with either sign depending on which assignment rule dominated during the period examined. Nothing about matching on observed size fixes this, because the risk signal that triggered coverage is usually the same signal that predicts the outcome. It needs either an actual randomised or staggered rollout of coverage, or a design built on a capacity constraint or territory boundary that assigns coverage for reasons unrelated to account health.

02

Treating raw request or usage volume as engagement

Most traffic in this domain is emitted by machines. Continuous-integration pipelines, scheduled batch jobs, synthetic monitors, backfills and client retries can all grow by an order of magnitude from one configuration change made by one engineer, and none of it represents a new decision to use the product. The inversion is what makes it dangerous: when the platform degrades, clients retry, so error-driven retry volume rises at the exact moment the customer is most likely to leave, and an engagement dashboard built on raw counts shows growth immediately before a churn. Filter on traffic_class and on successful status before anything else, and keep failed-request volume as its own separate series.

03

Reading experiment results before checking the arm split

Compare observed arm counts against the intended allocation ratio, not an assumed even split, and set the alarm far below the conventional 0.05: at 0.05 roughly one healthy experiment in twenty trips it, which is why sample-ratio checks usually run at p < 0.001 or stricter. The test's power scales with sample size, so it misses a real diversion on a small experiment and fires on an imbalance too small to move the estimate on a very large one. A flag means go find the assignment or logging fault before reading any outcome, not report a mismatch.

04

Never asking what decision the analysis will inform

Open with who makes the decision, what the options are, and by when. The answer determines the precision you need, the segments worth cutting, and whether an observational read suffices or an experiment is required.

Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.

15 technical prompts3 include a worked solution

How do you test for normality, and what non-parametric alternatives do…

medium
statistics and probability

How do you test for normality, and what non-parametric alternatives do you use when your data violates standard distribution assumptions?

Approach
  1. Write down the assumption the method needs before you use the method.
  2. Quantify uncertainty explicitly rather than reporting a point estimate alone.
  3. Say what the estimate is of, and over what population it generalises.
Follow-up
  • Which assumption here is most likely to be violated in practice?
  • How would you explain this result to someone who does not know statistics?

How do you monitor model drift and data degradation after deploying a …

medium
machine learning and modelling

How do you monitor model drift and data degradation after deploying a machine learning model to production?

Approach
  1. Say how the offline result would be validated online before it is trusted.
  2. Pick an evaluation metric that matches the cost of each error type, not a default.
  3. Set a baseline first, so any model has something honest to beat.
Follow-up
  • How would you choose the decision threshold, and who owns that choice?
  • Where could label leakage enter this setup?

Sessionise an API event stream with a 30-minute inactivity gap

mediumWorked solution
sessionisationevent-streamspandas

fct_api_request arrives as a DataFrame with account_id, user_id, request_at (tz-aware UTC), traffic_class and http_status, roughly 5 million rows. Assign a session_id to every human-attributable request: drop rows where user_id is null or traffic_class is in ('ci','synthetic_monitor','load_test'), then open a new session whenever the gap since that user's previous remaining request exceeds 30 minutes. Return the filtered frame plus session_id, and a per-session summary with user_id, account_id, session start, session end and request count. Do not loop over rows.

Approach
  1. Settle the filter-then-gap ordering before writing code. Removing CI and synthetic rows changes the gaps, so sessionising the raw stream and filtering afterwards is a different answer; the definition given filters first, and the two diverge most for accounts whose CI runs every ten minutes.
  2. Sort once by (user_id, request_at) with a stable kind, then gap = df.groupby('user_id', sort=False).request_at.diff(). The first row of each user yields NaT, which is exactly the boundary condition you want rather than a special case to patch.
  3. new_session = gap.isna() | (gap > Timedelta(minutes=30)); session_id = new_session.cumsum(). The cumsum runs over the whole sorted frame and therefore produces globally unique ids in one pass; a per-user cumcount collides across users and forces a composite key on every downstream join.
  4. Build the summary with a single groupby('session_id').agg(...). user_id and account_id can be carried with 'first' only because the sort key groups them — state that dependency, since it silently breaks if someone later re-sorts the frame.
  5. Decide explicitly what a session means when one user_id holds memberships in several accounts: either add account_id to the sort and group keys, or document that sessions may cross accounts. Leaving it undecided produces sessions whose account_id is whichever row sorted first.
Worked solution 30 min
  1. human = df[df.user_id.notna() & ~df.traffic_class.isin(['ci','synthetic_monitor','load_test'])].sort_values(['user_id','request_at'], kind='mergesort').reset_index(drop=True)
  2. gap = human.groupby('user_id', sort=False).request_at.diff(); human['session_id'] = (gap.isna() | (gap > pd.Timedelta(minutes=30))).cumsum()
  3. summary = human.groupby('session_id').agg(user_id=('user_id','first'), account_id=('account_id','first'), start=('request_at','min'), end=('request_at','max'), n_requests=('request_at','size')).reset_index()
  4. assert summary.n_requests.sum() == len(human) and human.groupby('session_id').user_id.nunique().max() == 1
EXPECTED RESULThuman gains a dense session_id increasing in (user_id, request_at); summary has exactly one row per session, n_requests sums to len(human), and every within-session consecutive gap is at most 30 minutes while every same-user between-session gap exceeds it.
Follow-up
  • Where does 30 minutes come from, and how would you pick it from this data instead of from convention?
  • An engineer reused their personal key for a nightly batch job, so machine traffic carries a human user_id. How would you detect that, and should those requests form sessions?
  • How much does the session count change if you sessionise before dropping CI traffic rather than after?

Roughly 90 minutes a night on weekdays with one longer weekend block. The plan deliberately cuts scope rather than compressing everything, on the assumption that finishing one thing a night beats half-starting four.

Small steps. Visible outcomes.0 / 7 completed
ONE WEEK · YOUR PACE

Prepare, practise & reflect

One practical outcome each day. Spend longer where you need it.

0 / 7 done
01Fix the scope and set a baseline
  • Read the role description and write the three things the loop will almost certainly test, then write an explicit not-doing list for everything else and keep it visible all week.
  • Take one 20-minute SQL prompt and one 10-minute metric question cold, and write the single sentence that says what blocked each attempt, since that sentence is what decides which two topics get the most evenings.
  • Set the week's one rule: one problem finished to completion every night, including the night you only have 40 minutes.

Deliverable: A one-page scope with an explicit not-doing list and two cold attempts, each carrying one sentence on what blocked it.

Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗
02One query pattern, written three times
  • Choose the single pattern most likely to appear (a cohort retention grid, or a funnel counted by user) and write it three times from a blank file rather than editing the previous attempt.
  • On the third attempt, write the grain of every CTE as a comment before writing its body.
  • Stop at 90 minutes even if the third version is imperfect, and write the one thing you would fix with another hour.

Deliverable: Three independent versions of the same query plus a note on what changed between them.

Practice prompt ↗Practice prompt ↗Practice prompt ↗
03Only the statistics you will be asked to defend
  • Write, in under 200 words, how you would decide whether a difference between two groups is real: the test, its assumptions, and what you would switch to when an assumption fails.
  • Compute a 95 percent confidence interval for a difference in proportions by hand on realistic numbers, then write in one sentence what changes if the two samples are paired rather than independent.
  • Write your answer to "what does a p-value mean", check it against a definition, and delete the version that describes it as the probability the hypothesis is true.

Deliverable: A 200-word written answer and one hand-computed interval you can reproduce under pressure.

Practice prompt ↗Practice prompt ↗Practice prompt ↗
04One case, and the assumptions holding it up
  • Answer one product case aloud in 20 minutes with a recording running, then listen back with a pen and mark every claim you asserted without saying what it rested on: an assumed user behaviour, an assumed data source, an assumed baseline rate, an assumed grain.
  • Pick the three assumptions the recommendation actually depends on, write how you would check each one against data, and say which one being wrong would flip the recommendation rather than merely weaken it.
  • Write the four-step structure you used onto a card small enough to hold in working memory when you are nervous.

Deliverable: One recording, three load-bearing assumptions each with a written check, and a four-step structure card.

Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗
05Your own work, timed
  • Write a 90-second version and a four-minute version of your main project, and time both out loud rather than reading them.
  • Prepare answers to the two follow-ups that always come: what you would do differently, and how you knew it worked.
  • Put one number in the first sentence and be able to say exactly where that number came from and what it excludes.

Deliverable: Two timed narratives with one defensible number in the opening line.

Practice prompt ↗Practice prompt ↗
06The one full rehearsal, in a longer weekend block
  • Run a 60-minute mock covering query work, a case and a behavioural question in a single sitting with no breaks, because sustained attention is the thing evenings have not trained.
  • Immediately afterwards, and before hearing any feedback, write the three moments you lost the thread.
  • Spend the rest of the block only on those three moments, and on nothing you merely feel shaky about.

Deliverable: Mock notes naming three failure moments with a specific fix written under each.

Practice prompt ↗Practice prompt ↗
07Taper
  • Write the 20-minute warm-up you will actually do on the morning of the interview: one query you can already write from a blank file, one metric you can define out loud, and nothing you have never seen before.
  • Re-read only your own notes from this week, and open no new material.
  • Write down the logistics: the tool you will be asked to work in, whether lookups are allowed, and the sentence you will use when you do not know something.

Deliverable: A one-page card holding the case structure, the project numbers, and the logistics.

Practice prompt ↗Practice prompt ↗Worked solution ↗

Expand any day for tasks and deliverables. Your progress is saved on this device.

An answer without a quantity is hard to interrogate, so interviewers keep probing until they find one. Come with the baseline, the change, the window it was measured over, and how confident you were. If the effect never got measured, say so and say what you would have measured. Fabricated precision is worse than an honest gap.

Tell me about a time when you had to influence a cross-functional prod…

medium
behavioural and stakeholder questions

Tell me about a time when you had to influence a cross-functional product roadmap using data, despite pushback from stakeholders.

Approach
  1. State the situation in two sentences and spend the rest on your reasoning.
  2. Close with what you would do differently, concretely.
  3. Pick a story where you drove the decision, not one where you observed it.
Follow-up
  • What would you do differently if you ran that project again?
  • How did you know the outcome was caused by your change?

Announce a metric fix that cuts the headline number

medium
metric definitionsstakeholderscommunication

Weekly active organisations, the count on the company dashboard, has never excluded rows where dim_account.is_internal is true, and it counts traffic with traffic_class in synthetic_monitor and load_test. Correcting both reduces that count by 11 percent and removes most of the growth reported over two quarters. The figure appears in a board deck and in two teams' quarterly goals, one written on the count and one on the weekly active organisation ratio, whose denominator is accounts whose account_status was in ('trial','free','active_paid') through the week. Decide the order in which you tell people, what the dashboard shows during the transition, and what you propose happens to goals already set against the old definition.

Approach
  1. The interviewer is probing whether you can land a correction as an operational change with a plan attached, rather than as an announcement other people then have to clean up after.
  2. Quantify each exclusion separately before telling anyone: internal accounts, synthetic monitors, load tests. Three known quantities are a discussion; one alarming total is an argument.
  3. Be precise about which side of the metric each exclusion touches, because one team's goal is on a count and the other's is on a ratio. The traffic-class filters remove requests, so they shrink the numerator only. Dropping internal accounts removes them from the ratio's denominator as well, since internal accounts carry ordinary account_status values and therefore sit in that denominator. Internal accounts are active in almost every week while the real base is not, so the numerator loses a larger share than the denominator and the ratio falls by less than the count does. Compute both and say which one the 11 percent is before anybody assumes.
  4. Check whether the trend changes, not only the level. A constant 11 percent shift is a rebasing and nothing more. A shift that widens over time means the reported growth was partly internal or synthetic, which makes the existing goals unachievable as written and changes what you are asking teams to do.
  5. Sequence the disclosure: the metric owner and the two teams whose goals move first and privately, then the board channel with a written bridge, then the dashboard. The dashboard is last because a number that changes without explanation is read as instability rather than as a fix.
  6. Run both series for one reporting period with the bridge visible, restate history rather than letting the series break at a date, and set the date the old series is removed.
  7. Propose the goal treatment yourself: rebase each target by the shift measured on the metric that target is written against, rather than leaving each team to negotiate individually, which is where corrections of this kind usually die.
Follow-up
  • One team's quarterly goal is now unreachable. Rebase the target or let it miss, and what does each choice teach the organisation?
  • How would this have been caught when the metric was first defined?
  • What else on that dashboard shares this failure mode, and how would you find out this week?

Report an underpowered consumption test to a non-technical executive

medium
communicationuncertaintyexperimentation

An account-randomised packaging change ran six weeks across 900 paying accounts. The effect on billable units per account per month is plus 4.1 percent, with a 95 percent interval from minus 3.2 to plus 11.8 after clustering standard errors at the account and applying the pre-registered winsorisation at the 99th percentile. An executive with no statistical background wants one number this week to decide a full rollout. Produce a three-sentence spoken answer, one chart, and an explicit recommendation of ship, stop or keep running, with the cost of each option stated.

Approach
  1. The interviewer is probing whether you can be decision-useful without either hiding the uncertainty or hiding behind it. Start from the decision rather than the statistics: establish what the executive would do differently at plus 4 percent versus zero, because if the action is identical the interval does not matter.
  2. Translate the interval into consequences in units the executive already reasons about. Multiply both endpoints by the cohort's baseline consumption and contracted rates to give an annualised revenue range, so the answer is a range of dollars rather than a range of percentages.
  3. Price the option to wait. Using the observed variance, state roughly how many additional account-weeks halve the interval width, so keep running becomes a quantified choice instead of a stall.
  4. Offer a cheaper path to the same decision: a lower-variance proximate outcome such as successful billable units on the new SKU, or CUPED using each account's pre-period consumption, quoting the expected variance reduction as one minus the squared pre-post correlation.
  5. Give a recommendation and name the single observation that would reverse it. A strong answer commits; a generic one recites the interval and leaves the decision on the table.
Follow-up
  • The executive says it clearly works and is just not provable, so ship it. What is your answer?
  • How much of the interval width comes from clustering and how much from the revenue tail, and what would you do about each?
  • If you had to ship this week with no more data, which guardrail would you watch for the first fortnight and at what threshold would you roll back?
  • 01

    Tell me about a time when you had to influence a cross-functional product roadmap using data, despite pushback from stakeholders.

  • 02

    Weekly active organisations, the count on the company dashboard, has never excluded rows where dim_account.is_internal is true, and it counts traffic with traffic_class in synthetic_monitor and load_test. Correcting both reduces that count by 11 percent and removes most of the growth reported over two quarters. The figure appears in a board deck and in two teams' quarterly goals, one written on the count and one on the weekly active organisation ratio, whose denominator is accounts whose account_status was in ('trial','free','active_paid') through the week. Decide the order in which you tell people, what the dashboard shows during the transition, and what you propose happens to goals already set against the old definition.

  • 03

    An account-randomised packaging change ran six weeks across 900 paying accounts. The effect on billable units per account per month is plus 4.1 percent, with a 95 percent interval from minus 3.2 to plus 11.8 after clustering standard errors at the account and applying the pre-registered winsorisation at the 99th percentile. An executive with no statistical background wants one number this week to decide a full rollout. Produce a three-sentence spoken answer, one chart, and an explicit recommendation of ship, stop or keep running, with the cost of each option stated.

PracHub interview preparation framework ↗
Is this an official Atlassian interview guide?

No. It is PracHub's own research and practice material for the Data Scientist role at Atlassian. Rounds and questions reflect what candidates have reported, not a process Atlassian has published, and they change over time. Confirm the current format and scope with your recruiter.

PracHub interview research ↗
How difficult is the interview process, and how much preparation time should I expect?

The interview loop is moderately to highly rigorous, emphasizing both technical depth and practical product judgment. Most candidates spend between four to six weeks in intensive preparation, focusing heavily on advanced SQL, experimentation design, and structured case studies.

PracHub interview research ↗
What differentiates successful candidates from those who do not pass?

Successful candidates stand out by structuring ambiguous problems methodically before diving into math or code. They combine technical fluency with strong business acumen, demonstrating how their analytical recommendations directly drive product strategy and user value.

PracHub interview research ↗
How should I approach the behavioral and values interview rounds?

Interviewers look for concrete examples of ownership, collaboration, and resilience, especially when navigating disagreement or project ambiguity. Ground your answers in past professional experiences, clearly highlighting your specific contributions and the ultimate business impact of your work.

PracHub interview research ↗
Are interviews conducted remotely, and how does the distributed model affect the process?

All interview rounds are conducted virtually via video conferencing and shared online coding environments. Because the company operates as a distributed-first organization, demonstrating strong asynchronous communication skills and comfort with virtual collaboration is a distinct advantage.

PracHub interview research ↗
Sources & methodology 3 sources ↗

Official role evidence, timestamped platform data and clearly labeled preparation advice.