1 digit technology Pvt · Data Scientist
Updated · 2026-10-02

1 digit technology Pvt Data Scientist
Interview Questions & Guide 2026

THE 60-SECOND BRIEF

Candidates describe 1 digit technology Pvt as a company working in insurance and technology, where the Data Scientist role touches product features, user engagement and product metrics. One reported question asks how to measure the success of a new insurance product feature, so prepare product cases set in an insurance context. Little else about the products is reported, so use the recruiter call to find out which product the team supports. As candidates describe the role, you design experiments, build predictive models, agree with product managers on what success means for a feature, diagnose performance problems, tune classification or regression pipelines, explain findings to non-technical colleagues, and advise engineering teams on what data to collect.

This guide is for candidates interviewing for the Data Scientist role at 1 digit technology Pvt, whether fresher or experienced. It covers the four reported rounds, the reported product-sense, SQL, A/B testing, machine learning and behavioral questions, and a 7-day plan that ends in a full mock. Candidates report that freshers should lean on logical reasoning and mathematical foundations, while experienced hires should show influence on product outcomes. Python, SQL, scikit-learn, TensorFlow or PyTorch, and A/B testing are the skills listed for the role.

Candidates report four rounds over roughly 3-5 weeks: Initial Screening, Technical Assessments, Behavioral Assessments and Stakeholder Interaction. Reports say some stages may be combined or sped up depending on the team, and mention both remote assessments and in-office meetings.

Decompose a metric move by segment and mixTurn a vague request into a measurable questionPick a randomisation unit that respects interference

45 min read

Practice 16 Data Scientist prompts
16Practice promptsAcross five skill areas
3With worked solutionsIncluded in the practice prompts

The reported question set mixes five kinds of work: diagnosing and designing product metrics, SQL with window functions, A/B testing and statistics, machine learning from classification pipelines to NLP, and behavioral prompts about methodology and disagreement. Because the role is described as sitting between data and product strategy, the product-metric questions are not a side topic. Several reported questions are phrased as business situations, such as a sudden drop in a core metric or the success measure for an insurance feature, so start your answer from the decision the metric has to support.

For SQL, the reported examples are the third highest salary in a table, running totals, cohort analysis, and, among other reported examples, the top 3 users by spend per month. All of them come down to choosing the right window function and frame: DENSE_RANK versus ROW_NUMBER when values tie, a frame with a unique tiebreaker in ORDER BY for a running sum, and a first-activity date computed per user to assign cohorts. For experimentation, be ready for sample-size calculation, one-tailed versus two-tailed tests, seasonality and outliers. Statistical significance, p-values and confidence intervals are listed as foundations, and one reported prompt asks you to explain significance to a non-technical stakeholder.

The machine learning side is broad. Reported material names NLP, deep learning and Python-based data manipulation, a multi-class classification problem from start to finish, traditional models versus Transformer-based architectures for NLP, and the practical use of Open CV in past projects. Scikit-learn, TensorFlow and PyTorch appear in the listed skills. Prepare your own past projects in enough detail to say which library you used, why, and what the result was in business terms. Candidates are also told to be ready to quantify the impact of resume projects.

Reported tips add a few habits worth rehearsing. Use STAR for behavioral answers. Open a case by asking what the business wants from the analysis, agree which metrics will show it, and list the unusual cases that could distort them before you choose a technique. Expect your assumptions to be questioned and treat that as a joint working session rather than an attack. For logic puzzles, candidates are told the reasoning steps matter more than the mechanics of the puzzle, so practise thinking aloud.

01

Initial Screening

reported

Candidates describe the Initial Screening as the opening stage, where fit for the role is assessed. Little more is reported about its format, so prepare a short summary of your background and map yourself against the listed skills: Python, advanced SQL, a machine learning framework such as scikit-learn, TensorFlow or PyTorch, and A/B testing. Reports also suggest that freshers lean on logical reasoning and mathematical foundations, while experienced hires lean on influence on product outcomes. That advice is given for the role overall, not for this stage, but it tells you how to frame your background. Choose which framing is yours and have one project ready where you can state a measured business result.

What to demonstrate

  • Fit for the role, which is how candidates describe the purpose of this stage
  • How clearly you can summarise your background and your reasons for wanting a product-focused data science role, which a screening conversation inherently covers

How to prepare

  • Write a short spoken summary of your path that names the stack you have used and ends with one project result expressed as a number
  • Mark each listed skill as used in production, used in coursework or a side project, or not used, and have a one-sentence answer for each gap
  • Pick your framing: freshers lean on reasoning and maths, experienced hires on influence over product outcomes, and prepare one project that supports it
  • Ask the recruiter about the format of the later rounds, including which parts are remote and which are in person
PracHub interview research ↗
02

Technical Assessments

reported

Candidates report that the Technical Assessments cover logical reasoning and machine learning. Across the loop, reports mention both remote assessments and, in many cases, in-office meetings, so ask the recruiter which applies to this stage. Candidates also list SQL, statistical knowledge and deep learning among the topics tested without tying them to one round, so keep all of them warm for this stage. Prepare to write working code and to state your reasoning as you go. For logic questions, name each step you take and avoid getting lost in puzzle mechanics. For the machine learning topics, practise a full pass from problem framing through model choice to evaluation.

What to demonstrate

  • Logical reasoning, shown through the steps you state while solving
  • Machine learning judgement: framing, model choice and evaluation for problems such as multi-class classification
  • If the assessment includes SQL, Python or statistics (all listed for the role but not tied to a round), code and queries that return the right result and statistical reasoning you can explain

How to prepare

  • Build a multi-class classifier in scikit-learn end to end: baseline, pipeline, cross-validation, macro-averaged metric and confusion matrix
  • Write a one-page comparison of TF-IDF with a linear model versus a fine-tuned Transformer for a text classification task, covering data size, latency, cost and interpretability
  • Keep SQL warm even though no report ties it to this stage: write a ranking query and a running total from a blank file and test both on tied values
  • Solve two logic puzzles aloud, naming each deduction before you make it
PracHub interview research ↗
03

Behavioral Assessments

reported

Candidates describe the Behavioral Assessments as the stage that evaluates soft skills and cultural fit. The reported behavioral questions, which are not tied to a specific round, include defending a technical methodology to a skeptical stakeholder, pivoting after unexpected data findings, disagreeing with colleagues over how to read model results, and mentoring a colleague or contributing to a team's growth. Reports also advise preparing to discuss how you take feedback and navigate disagreement, and to use STAR. Prepare a distinct story for each prompt so each one shows a different decision of yours.

What to demonstrate

  • Soft skills and cultural fit, which is how candidates describe this round
  • How you handle feedback and navigate disagreement, which reports advise preparing to discuss
  • How you contribute to a team, the other area reports name for this kind of discussion

How to prepare

  • Write four STAR stories, one for each reported prompt, and give each a result with a number and its source
  • For the methodology-defence story, write down the specific objection the stakeholder raised and the evidence that answered it
  • For the pivot story, state what the unexpected finding was, what you changed in the approach and what happened afterwards
  • Rehearse each story aloud: keep the situation to two sentences and spend the rest on what you decided and did yourself, as distinct from what the team did
PracHub interview research ↗
04

Stakeholder Interaction

reported

Candidates report that Stakeholder Interaction puts you in front of various stakeholders, including lead data scientists and senior leadership. That mix suggests two registers: technical depth on your models and experiments for data science peers, and business framing for leadership. Clear communication of complex ideas is listed among the role's soft skills, so be ready to describe statistical significance or a model result to someone with no statistics background. Confirm with the recruiter who attends, because the preparation differs by audience.

What to demonstrate

  • The audiences reported for this stage: lead data scientists and senior leadership
  • Explaining complex technical work clearly, which candidates list among the role's soft skills

How to prepare

  • Prepare a short business version and a longer technical version of your best project, with the same numbers in both
  • Write a plain-language explanation of statistical significance and of a confidence interval, with no jargon
  • Re-read your earlier answers and list the figures you quoted so they match when repeated
  • Candidates list handling constructive criticism during technical deep dives among the role's soft skills, so decide how you will respond when someone challenges a core assumption: restate it, say what evidence would change your view, and continue
PracHub interview research ↗

PracHub editorial advice for the preparation topics above.

01

Answering a sudden metric drop by suggesting fixes before checking the number is real

Start with validation: check for instrumentation or logging changes, pipeline delays, and a change in the metric definition. Then slice by platform, region, acquisition channel, new versus returning users and product surface to localise the drop, and only then form causes such as a release, a campaign or seasonality. Say aloud what each check would rule out.

02

Picking a ranking method for the third highest salary without settling how ties count, then returning duplicates or an empty result

Ask whether tied salaries count once. If they do, use DENSE_RANK() or SELECT DISTINCT salary ... ORDER BY salary DESC LIMIT 1 OFFSET 2. If each employee counts separately, use ROW_NUMBER() (or RANK() if ties should share a place) or LIMIT/OFFSET without DISTINCT. Say what the query returns when fewer than three values exist (NULL or no row). For running totals, add a unique tiebreaker to ORDER BY (for example order_date, order_id) and then use ROWS, or keep RANGE if all rows on a date should share one cumulative value.

03

Choosing a one-tailed test after seeing which direction the result went, or quoting a p-value as the chance the hypothesis is true

Decide tail, significance level, primary metric and minimum detectable effect before launch. Use two-tailed unless a loss in the opposite direction truly would not change the decision. Explain a p-value as the probability of data at least this extreme if there were no effect, and give a confidence interval alongside it.

04

Recommending a Transformer for an NLP task by default, with no baseline or cost reasoning

Start from the constraints: labelled data volume, latency, serving cost, interpretability and how much the task depends on word order and context. Name a TF-IDF plus linear model baseline, compare it on a held-out set, and move to a pretrained Transformer only when the measured gain justifies the cost.

05

Telling project and behavioral stories in terms of what the team did, with no personal decision and no quantified result

Use STAR with 'I' for your own actions. Put one number in the result, know where it came from and what it excludes, and make sure the same figures appear each time you retell the story.

Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.

13 technical prompts3 include a worked solution

What are the trade-offs between using traditional ML models versus Tra…

medium
machine learning and modelling

What are the trade-offs between using traditional ML models versus Transformer-based architectures for NLP tasks?

Approach
  1. Start with the traditional baseline: TF-IDF features over word or character n-grams fed into logistic regression, a linear SVM or Naive Bayes. They train in minutes on CPU, serve with very low latency, have inspectable coefficients, and are often strong on keyword- or topic-driven tasks with limited labels.
  2. Transformer models (fine-tuned BERT-family encoders, or prompting a large language model) use contextual representations. They handle word order, negation, synonyms and longer context much better, and pretraining transfers knowledge, so they often reach higher accuracy with fewer labels on semantic tasks.
  3. Name the costs: training usually needs GPUs, and self-attention cost grows quadratically with sequence length, which raises inference latency and serving cost. Encoders such as BERT have a fixed maximum input length (512 tokens), so long documents need truncation or chunking. Predictions are harder to explain, and models are larger to deploy and retrain.
  4. Decide on concrete criteria: label volume, how semantic the task is, the latency and throughput budget, the explainability requirement (for example, if a decision has to be justified to a customer), and the maintenance cost. Middle-ground options include frozen sentence embeddings with a linear head, or a distilled smaller transformer.
  5. Settle the choice empirically: compare both on the same held-out split using the metric that matters for the business, do error analysis on the cases where the transformer wins, and check whether that gain justifies the extra serving and maintenance cost.
Follow-up
  • You have only 500 labelled examples. What do you try first, and why?
  • How would you serve a transformer model under a tight latency budget?
  • How would you explain an individual transformer prediction to a non-technical stakeholder?

Describe a situation where you had to clean a messy dataset to get it …

medium
machine learning and modelling

Describe a situation where you had to clean a messy dataset to get it model-ready.

Approach
  1. Pick one real project and name the concrete defects, not just 'messy data': missing values by column, duplicate records, inconsistent category spellings, mixed types or units, timezone mismatches, impossible values. Say how you found them: profiling in pandas with null rates, value counts, range checks and key-uniqueness checks.
  2. Explain each decision and its reason. Deduplicate on a defined business key. Treat missing values according to why they are missing: add a missing-indicator feature when missingness may carry signal, and impute with statistics computed on training data only. Separate outliers that are data errors (fix or drop) from real extreme values (keep, cap or transform).
  3. Show leakage awareness: fit imputers, scalers and encoders inside a scikit-learn Pipeline or ColumnTransformer so cross-validation never sees statistics from held-out folds, and drop fields recorded or updated after the prediction moment.
  4. Make the cleaning reproducible and reusable at prediction time: a scripted, versioned step with assertions on schema and value ranges, applied identically in training and serving, rather than one-off notebook edits.
  5. Quantify the outcome: how many rows were dropped or repaired, how the class balance changed, and the measured effect on the validation metric compared with training on the uncleaned data. State the business consequence in one sentence.
Follow-up
  • How did you decide between dropping rows and imputing them, and how did you check that dropping did not bias the sample?
  • How did you make sure the same cleaning ran on new data at prediction time?
  • What would you do differently if the dataset were ten times larger?

Rebuild per-visitor ordering without groupby convenience methods

easyWorked solution
pandasvectorisationwindow logic

You have a DataFrame of 2 million fct_event rows with visitor_id, occurred_at_utc and event_id, unsorted and containing duplicate timestamps within a visitor. Produce three new columns: event_rank, the 1-based position of the event within its visitor ordered by occurred_at_utc; seconds_since_prev, the gap to that visitor's previous event, NULL for the first; and is_first_for_visitor. You may use sort_values, shift, cumsum, numpy and boolean masking. You may not use groupby.transform, groupby.apply, groupby.cumcount, groupby.rank or merge_asof. Break timestamp ties on event_id.

Approach
  1. Sort once by ['visitor_id', 'occurred_at_utc', 'event_id'] and reset the index. The whole exercise reduces to row arithmetic on a sorted frame, and the tiebreak on event_id is what makes the result reproducible across runs.
  2. Mark visitor boundaries with is_first = df['visitor_id'].ne(df['visitor_id'].shift()). This is the single fact every other column derives from.
  3. Compute seconds_since_prev as the diff of the timestamp column, then overwrite it with NaT/NaN wherever is_first is True. The shift crosses the boundary between visitors and will otherwise hand the first row of each visitor the last event of the previous one.
  4. Build event_rank from a running counter that resets at boundaries: take a global cumulative position (np.arange(len(df))) and subtract, per row, the global position at which that visitor started. Get the start position by forward-filling the positions where is_first is True, which is a cumsum-free reset and is O(n).
  5. Verify against the forbidden method once, as a test rather than as the implementation, and confirm the two agree on every row.
Worked solution 20 min
  1. Sort on the three-key tuple and reset_index(drop=True).
  2. Compute is_first via .ne(.shift()), which is True for row 0 because the shifted value is NaN.
  3. pos = np.arange(len(df)); start = pd.Series(np.where(is_first, pos, np.nan)).ffill(); event_rank = (pos - start + 1).astype(int).
  4. gap = df['occurred_at_utc'].diff().dt.total_seconds(); gap[is_first] = np.nan.
  5. Assert event_rank equals df.groupby('visitor_id').cumcount() + 1 on the sorted frame.
EXPECTED RESULTThree columns on the sorted frame: event_rank starting at 1 for every visitor and increasing by 1 with no gaps, seconds_since_prev null exactly where is_first_for_visitor is True, and is_first_for_visitor summing to df['visitor_id'].nunique().
Follow-up
  • The frame does not fit in memory. How does your approach change if you can only process one visitor-partitioned chunk at a time?
  • occurred_at_utc is client-supplied and sometimes runs backwards within a visitor. Does your seconds_since_prev go negative, and should it?
  • How would you extend this to reset the counter at every change of surface as well as visitor?

The plan follows the four reported rounds and the reported question types: product metrics first, then SQL, then experimentation and statistics, then machine learning, then data-collection trade-offs, then behavioral stories, ending with a full rehearsal. Each day produces something you can show.

Small steps. Visible outcomes.0 / 7 completed
ONE WEEK · YOUR PACE

Prepare, practise & reflect

One practical outcome each day. Spend longer where you need it.

0 / 7 done
01Map the loop and write two cold product answers
  • List the four reported rounds and the skills named for the role (Python, SQL, scikit-learn, TensorFlow or PyTorch, A/B testing), and mark each as strong, rusty or new
  • Write a cold outline for the question about a sudden drop in a core product metric, as a sequence of checks
  • Write a cold outline for designing a metric to measure a new insurance product feature, naming a primary metric, a guardrail and the decision it supports
  • Write three project impact lines, each with one number and a note on where it came from

Deliverable: A one-page skills map, two cold outlines for the product questions, and three quantified project lines.

Practice prompt ↗Practice prompt ↗Worked solution ↗
02SQL window functions
  • Write the third-highest-salary query three ways (DENSE_RANK, a DISTINCT subquery, OFFSET without DISTINCT) and test each with tied salaries and with fewer than three distinct values
  • Write a running total with SUM() OVER (PARTITION BY ... ORDER BY order_date, order_id ROWS ...), then compare it with RANGE on order_date alone and note how tied dates differ
  • Write a cohort retention query that assigns each user a cohort from their first activity date, then counts active users per cohort and month
  • Write the top 3 users by spend per month using a ranking function partitioned by month, and complete the reactivation-gap practice problem

Deliverable: Three third-highest-salary variants tested on ties and short tables, a running total with a ROWS versus RANGE comparison, a cohort retention query, a top-3-per-month query, and the reactivation-gap practice solution.

Practice prompt ↗Practice prompt ↗Practice prompt ↗
03Experimentation, statistics and logic puzzles
  • Work out the sample size per arm for a conversion experiment from a baseline rate and a minimum detectable effect, first by formula and then with a library call, and check they agree
  • Write when a one-tailed test is defensible for a launch decision and why a two-tailed test is the safer default, then write a plain-language explanation of statistical significance for a non-technical stakeholder
  • List the pitfalls you have met in your own projects (peeking, sample-ratio mismatch, novelty effects, multiple comparisons, seasonality, outliers), write how you detected or handled each, then complete the shared-workspace randomisation practice problem
  • Solve two logic puzzles aloud, naming each deduction before you make it, and do not get pulled into the puzzle's mechanics

Deliverable: A sample-size worksheet, a one-tail versus two-tail note with a plain-language significance explanation, a pitfalls list with a detection method for each plus the randomisation practice solution, and notes from two spoken puzzle solutions.

Practice prompt ↗Practice prompt ↗Practice prompt ↗
04Machine learning end to end
  • Build a multi-class classifier in scikit-learn on a public dataset: baseline, preprocessing pipeline, cross-validation, macro-averaged metric and confusion matrix
  • Using pandas, clean a messy public dataset (null rates, duplicates, value counts, a groupby, a merge, per-user ordering), write what you repaired, the leakage risks and what you checked afterwards, then complete the per-visitor ordering practice problem
  • Fine-tune a small pretrained text model in PyTorch or TensorFlow on a public text-classification dataset and compare it with a TF-IDF plus logistic regression baseline on accuracy, training time, inference latency and interpretability
  • Prepare a two-sentence account of how you used Open CV in a past project, or of the closest computer vision work you have done

Deliverable: A working classifier notebook, a pandas cleaning notebook with a written account, the per-visitor ordering practice solution, a fine-tuned text model with a comparison table against the TF-IDF baseline, and the Open CV answer.

Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗
05Metric diagnosis and data-collection trade-offs
  • Rehearse the drop answer for a core metric that falls 10% overnight: list the first three checks and what each would rule out
  • Write your answer to balancing user experience with business-driven data collection, covering what you would collect, what you would drop and who decides
  • Work the practice problem where a pooled conversion rate falls while every segment rises, and write the mix-shift explanation in two sentences
  • Record yourself answering one product case aloud and note each assertion that has no stated basis

Deliverable: A three-check metric-drop script, a data-collection trade-off answer, the mix-shift explanation and a recording with annotated assumptions.

Practice prompt ↗Practice prompt ↗Practice prompt ↗
06Behavioral stories
  • Write one STAR story each for defending your methodology to a skeptical stakeholder, pivoting after an unexpected data finding, and mentoring or growing a team
  • Write the disagreement story about interpreting model results: both positions, the agreed check, what it showed and what you changed
  • Practise the impact-ownership and flat-experiment practice problems so you can state your own contribution without claiming the topline
  • Say each story aloud, shorten the setup until it takes two sentences, and list every figure you quote so you repeat them consistently

Deliverable: Four STAR stories with a result number and its source, written answers to the two practice problems, and a list of the figures you quote.

Practice prompt ↗Practice prompt ↗Practice prompt ↗
07Full mock and taper
  • Run a mixed mock: talk through one logic puzzle aloud, write one SQL query, answer one experiment question and one machine learning question
  • Finish with a stakeholder-style question: explain one result to a listener with no statistics background
  • Write a single page holding your case structure (objective, metrics, edge cases), your project numbers and the questions you will ask the recruiter
  • Read only your own notes from the week and open no new material

Deliverable: A mock log with three weak moments and a fix for each, plus a one-page review sheet.

Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗

Expand any day for tasks and deliverables. Your progress is saved on this device.

Candidates describe the Behavioral Assessments as the stage covering soft skills and cultural fit, and reports advise being ready to talk about feedback, disagreement and contributing to a team. The reported prompts cover defending a methodology, changing direction when the data surprises you, disagreeing over what model results mean, and mentoring. For a data scientist, the strongest stories show your reasoning and your own decisions, end in a measurable result, and show you can explain technical choices to a non-technical person.

How do you handle disagreements with team members regarding the interp…

medium
behavioural and stakeholder questions

How do you handle disagreements with team members regarding the interpretation of model results?

Approach
  1. Pick a disagreement where you owned part of the outcome and the subject was how to read results, for example whether a lift was real, which metric should decide, or whether a model's errors were acceptable for a segment. Avoid a story where you were simply right and the colleague was uninformed.
  2. Describe both positions in their strongest form in a sentence each, so the listener sees you understood the other reading. For instance: I thought the offline AUC gain would not survive deployment, and my colleague thought it would.
  3. Move quickly to how the disagreement was settled with evidence. Name the specific check you both agreed to, such as a holdout comparison, a calibration plot, an error breakdown by segment, a sensitivity analysis or a pre-agreed decision rule, and say what it showed.
  4. Say what you did that was personal: how you raised it, what you changed in your own analysis, and whether your position moved. Include a case where the other person was right, or partly right, because a story where you changed your mind shows you separate your ego from the result.
  5. Close with the outcome in business terms, any quantified effect, and what you do now to avoid the same disagreement, such as agreeing on the evaluation metric and decision threshold before training starts.
Follow-up
  • What would you have done if the data could not settle the disagreement?
  • How did you keep the working relationship intact afterwards?
  • Would you handle it differently if the other person were a senior stakeholder rather than a peer?

Quantify your own impact without claiming the topline you touched

hard
self-assessmentattributioncommunication

You are writing the impact section of your own review. Over the year you ran four experiments, one of which shipped and three of which were flat; you corrected the definition of gross monthly revenue churn so that cancellation is recognised at period_end_utc; and you built a self-serve funnel dashboard. Weekly active accounts rose 14% over the same period. Your reviewer knows the data well. Write the three impact claims you would defend, stating for each what you contributed, what evidence supports it, and what portion of the outcome you are not claiming.

Approach
  1. Recognise what is being probed: whether you apply to your own work the causal standard you would apply to somebody else's roadmap claim. Nearly everyone who would reject 'accounts that do Y retain better' will write 'I drove a 14% increase' without noticing it is the same error with a friendlier subject.
  2. Sort the work by the kind of evidence it can carry. The shipped experiment is the only item with a randomised estimate, so it is the only one where an effect size is defensible, and you claim the interval rather than the point estimate.
  3. Claim the three flat experiments as decisions prevented and price them. Features not built, or built differently, on evidence, with the engineering weeks reallocated as the number somebody else can verify. A defensible null is a delivered decision and should be written as one.
  4. Claim the definition fix as correctness, not as improvement. The old figure was overstated by a specific percentage and appeared in a specific set of recurring documents; the impact is the change it produced in the forecast built on top of it, not a change in churn itself.
  5. Claim the dashboard on usage and displacement: distinct weekly users of it, and the ad-hoc request count for six months before against six months after. If the request log does not exist, record the claim as unverified rather than estimating it upward.
  6. Disclaim the 14% explicitly and once. State that it cannot be separated from seasonality, other teams' launches and a pricing change, and bound your own contribution from above using the shipped experiment's interval converted into headline units.
Follow-up
  • Your shipped experiment's interval was +0.2pp to +1.4pp on activation. How much of the 14% can that account for, and how do you say so without undercutting yourself?
  • A peer in the same cycle claims the full 14%. What, if anything, do you do about it?
  • If you could only keep two of your three claims, which do you drop, and why that one?

Defend a flat experiment readout against a post-hoc segment

medium
experimentssegmentationpushback

A feature you evaluated is flat on seven-day activation: +0.05pp with a 95% interval of [-0.47pp, +0.57pp], from 61,000 exposed users per arm in fct_experiment_exposure joined to dim_user and fct_event. Baseline activation is 32%. The launch team asks you to drop every surface except mobile_web, where the point estimate is +1.1pp, and re-run. You have ten minutes in their planning meeting. Deliver a spoken position: what you will and will not do, and the decision you recommend.

Approach
  1. Recognise what is being probed: whether you hold a statistical position under social pressure without becoming either rigid or apologetic. A generic answer says the segment is not significant; a strong one separates the request into a question that is answerable (is the mobile_web number real?) and one that is not (can we ship on it?), and answers both.
  2. Price the multiplicity out loud. The slice was chosen after seeing the results, so its estimate is selected on favourable noise and is biased away from zero. With k independent looks at a nominal 5% level, the chance of at least one false positive is 1 - 0.95^k: 26% at six segments, 64% at twenty. Quote the k you actually inspected, not the k you reported.
  3. Use the arithmetic already in front of you. On the point estimates, a +1.1pp mobile_web effect combined with a pooled +0.05pp implies the remaining surfaces average negative in proportion to mobile_web's share of exposures. State that as a testable implication of their story rather than as a rebuttal of it.
  4. Ask the one question that settles the category: was mobile_web named in the analysis plan before launch? If it was, it is a planned comparison and gets a corrected reading. If it was not, it is a hypothesis, and the honest move is to size the test that would confirm it.
  5. Convert the refusal into a cost. Size a mobile_web-only confirmatory test at the claimed effect, state the weeks of mobile_web traffic it needs, and close with the recommendation: do not ship this as a lift, and note that the interval already rules out anything at or above +0.6pp, which is itself a useful input to the roadmap.
Follow-up
  • The confirmatory test you sized needs nine weeks of mobile_web traffic and the team has three. What do you recommend instead?
  • Suppose mobile_web was pre-registered. How does your reading change, and what correction do you apply?
  • Your interval excludes +0.6pp. Is that the same as saying the feature does nothing?
  • 01

    Tell me about a time you had to defend your technical methodology to a skeptical stakeholder.

  • 02

    Describe a challenging project where you had to pivot your approach due to unexpected data findings.

  • 03

    How do you handle disagreements with team members regarding the interpretation of model results?

  • 04

    Tell me about a time you mentored a colleague or contributed to a team's growth.

  • 05

    Tell me about a project where you could show the business impact of your work in numbers, and explain where the numbers came from.

  • 06

    Describe a time you received tough feedback on an analysis and what you changed afterwards.

PracHub interview preparation framework ↗
Is this an official 1 digit technology Pvt interview guide?

No. It is PracHub's own research and practice material for the Data Scientist role at 1 digit technology Pvt. Rounds and questions reflect what candidates have reported, not a process 1 digit technology Pvt has published, and they change over time. Confirm the current format and scope with your recruiter.

PracHub interview research ↗
How many rounds are reported, and what are they?

Candidates report four rounds over roughly 3-5 weeks: Initial Screening, Technical Assessments, Behavioral Assessments and Stakeholder Interaction. Reports say some stages may be combined or sped up depending on the team, and mention both remote assessments and, in many cases, in-office meetings.

PracHub Data Scientist practice ↗
What SQL should I practise?

Window functions and complex joins are the reported focus. Practise ranking with DENSE_RANK and ROW_NUMBER (the third highest salary), running totals with an explicit frame and a unique tiebreaker, cohort analysis based on a first-activity date, and top-N per group, such as the top 3 users by spend per month. Test each query with ties, missing rows and empty groups.

PracHub Data Scientist practice ↗
Which tools and topics should I know?

Candidates list Python, advanced SQL, scikit-learn and TensorFlow or PyTorch, along with A/B testing. Reported topic areas are SQL, statistical knowledge, logical reasoning, general machine learning and deep learning, and the questions touch NLP, Transformers and Open CV. Be ready to give a concrete example from your own work for each one you claim.

PracHub Data Scientist practice ↗
Do I need insurance domain knowledge?

A reported tip suggests researching the company's product line, especially in insurance, since it helps with product-sense questions. At minimum, think through what a success metric for a new insurance feature could be, and what could go wrong with it, such as a metric that improves because of who chooses to use the feature.

PracHub Data Scientist practice ↗
How should freshers and experienced candidates present themselves differently?

Candidates report that freshers should emphasise logical reasoning and mathematical foundations, while experienced hires should highlight how they influenced product outcomes. In both cases, expect to state the business impact of resume projects in quantified terms, so prepare numbers you can explain and source.

PracHub Data Scientist practice ↗
How should I approach a case study question?

Candidates are told to open with what the business wants from the analysis, then agree which numbers will show it and which unusual cases could distort them, before choosing a method. They are also told to expect their assumptions to be challenged, so state your assumptions explicitly and say which evidence would change your answer.

PracHub Data Scientist practice ↗
Sources & methodology 3 sources ↗

Official role evidence, timestamped platform data and clearly labeled preparation advice.