Tata Consultancy Services · Data Scientist
Updated · 2026-09-24

Tata Consultancy Services Data Scientist
Interview Questions & Guide 2026

THE 60-SECOND BRIEF

A Data Scientist at Tata Consultancy Services (TCS) functions at the intersection of advanced analytics, business strategy, and technical implementation. You will be responsible for translating complex business requirements into actionable data models, providing insights that drive decision-making for large-scale enterprise clients. Your work directly impacts how organizations optimize their operations, enhance customer experiences, and leverage predictive modeling to maintain a competitive edge.

Scope your preparation by the data you would actually touch, because the title will not tell you. A seat that lives in event logs and weekly readouts rewards fluency in aggregation and metric definitions; a seat that owns a model in production rewards fluency in train/serve skew, retraining cadence and drift monitoring. The fastest way to find out which one you are interviewing for is to ask what the team shipped last quarter and what it gets paged about.

PracHub has no confirmed round sequence for Tata Consultancy Services. Treat the sections below as preparation areas and confirm the format with your recruiter.

Segment margin by pricing model before comparingSeparate bookings, recognised revenue and collected cashReconstruct pipeline stage history from mutable rows

32 min read

Practice 14 Data Scientist prompts
14Practice promptsAcross five skill areas
3With worked solutionsIncluded in the practice prompts

A Data Scientist at Tata Consultancy Services (TCS) functions at the intersection of advanced analytics, business strategy, and technical implementation. You will be responsible for translating complex business requirements into actionable data models, providing insights that drive decision-making for large-scale enterprise clients. Your work directly impacts how organizations optimize their operations, enhance customer experiences, and leverage predictive modeling to maintain a competitive edge.

This role is critical to the Tata Consultancy Services mission of providing high-value IT solutions. You will engage in the full lifecycle of data science projects, from raw data extraction and exploratory analysis to the deployment of machine learning models. Whether you are working on supply chain optimization, churn prediction, or personalized customer analytics, you will be expected to balance technical rigor with the practical realities of product-driven business environments.

01

Preparation focus

editorial

No round sequence has been reported for this company, so work the categories below and confirm the format with your recruiter.

What to demonstrate

  • Breadth across SQL, experimentation and product reasoning
  • Ability to state assumptions before choosing a method

How to prepare

  • Drill the practice exercises below and time yourself
  • Prepare three quantified stories about decisions you drove
PracHub interview preparation framework ↗

PracHub editorial advice for the preparation topics above.

01

Pooling margin, realisation or overrun across pricing models

Fixed-fee margin falls with hours worked; uncapped time-and-materials margin rises with hours worked; retainer margin depends on neither. A quarter in which the firm sells more fixed-fee work will show a margin change caused entirely by mix, not by delivery performance, and the aggregate can move in the opposite direction to every individual pricing model. Always stratify by fct_engagement.pricing_model before comparing periods, and report the mix shift alongside the within-stratum change.

02

Trending utilisation or revenue on work_date without accounting for timesheet backfill

Time entries are created days to weeks after the work happens, and the backfill tail often runs two to six weeks. A dashboard keyed on work_date therefore shows the most recent weeks as a decline that reverses on every refresh. The fix is either to hold the reporting window back past the observed backfill tail (measure the tail with the timesheet submission lag metric rather than guessing) or to report an as-of-entered_at snapshot so the series is internally consistent, and to state which one you used.

03

Over-explaining the method and under-explaining the implication

Lead with the answer and what you would do about it, then give the approach when asked. Roughly one sentence of method per three of implication is the right ratio for a stakeholder-facing answer; the interviewer already knows what a regression is.

04

Reporting a p-value with no effect size or interval

Give the estimated difference with a confidence interval in the units the business cares about, then say whether that whole interval is worth acting on. A p-value only addresses whether you can rule out exactly zero; it says nothing about magnitude.

Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.

11 technical prompts3 include a worked solution

Cluster bootstrap a margin change across few, unequal accounts

hard
cluster bootstrapfew clustersmix shift

engagements has engagement_id, client_id, pricing_model, fees_usd, cost_usd and period in {pre, post} around a rate-card change. There are about 35 client_ids across roughly 400 engagements, and the top five clients carry a large share of fees. Estimate the pre-to-post change in value-weighted gross margin within pricing_model, and give a 95% interval by resampling whole client_ids with replacement, B = 2000. Write the bootstrap yourself; no resampling library. Report the account-weighted estimate alongside and explain any divergence.

Approach
  1. Write the statistic as a pure function of a dataframe first. Within each pricing_model, margin = (sum(fees) - sum(cost)) / sum(fees) per period; the headline is the fee-weighted average of the within-model differences using fixed pre-period weights, so a shift in the mix of work sold cannot masquerade as a change in delivery.
  2. Resample clusters, not rows: draw 35 client_ids with replacement and concatenate all engagements of each drawn client. A client drawn twice contributes its rows twice under distinct pseudo-ids, which is what preserves within-account correlation instead of averaging it away.
  3. Recompute the statistic on each replicate and take a percentile interval from the 2.5th and 97.5th quantiles. With concentrated fees the replicate distribution is skewed, so a symmetric point plus or minus 1.96 times a standard error is wrong in the tail that matters.
  4. Compute the account-weighted version (mean across clients of each client's margin change) next to the value-weighted one, and attribute the divergence to concentration rather than to noise; if they disagree in sign, that fact is the finding.
  5. Name the regime honestly: at roughly 35 clusters both cluster-robust standard errors and the pairs cluster bootstrap under-cover, so state the fix you would run next, which is a wild cluster bootstrap with Rademacher weights or CR2 with t(G-1) critical values.
Follow-up
  • Implement the wild cluster bootstrap and show at what cluster count its p-value separates from the naive one.
  • One client is 30% of fees. Show the interval with and without that account and say which one you would present, and to whom.
  • What pre-period check would make you willing to call this a causal effect of the rate card rather than a correlation?

Collapse time entries into contiguous staffing spells

mediumWorked solution
sessionisationgap and islandvectorisation

From approved delivery time entries (consultant_id, engagement_id, work_date, hours, charge_code, status, with engagement_id not null), build staffing spells. Per consultant and engagement, collapse weeks containing any logged hours into contiguous runs, where three or more consecutive zero-hour weeks end a spell. Output one row per spell with consultant_id, engagement_id, start_week, end_week, active_weeks, gap_weeks and total_hours. Consultants sit on several engagements at once, so spells from different engagements may overlap in time and must not be merged. No Python loop over rows.

Approach
  1. Aggregate to (consultant_id, engagement_id, week_start) with summed hours and keep only weeks with positive hours. The absent weeks are the signal, so materialising zeros here would destroy the thing you are detecting.
  2. Convert week_start into an integer week index, ((week_start - epoch_monday).dt.days // 7), so gap detection is integer subtraction rather than calendar arithmetic that breaks over month and year boundaries.
  3. Sort by [consultant_id, engagement_id, week_index], diff the week index inside each pair, and mark a spell start where the diff is null (first row of the pair) or greater than 3. A diff of 1 is adjacent weeks and a diff of 3 is two empty weeks, which the tolerance permits.
  4. Take spell_id = the cumulative sum of that boolean over the whole frame so ids are globally unique, then a single groupby on [consultant_id, engagement_id, spell_id] yields min and max week, active week count and summed hours; gap_weeks = (end - start + 1) - active_weeks.
  5. Do not deduplicate overlapping spells across engagements. A consultant on two engagements in the same week is the normal case, and that overlap is the fact any capacity or context-switching question needs.
Worked solution 30 min
  1. Filter to status == 'approved', charge_code == 'client_delivery' and engagement_id.notna(); derive week_start by subtracting the weekday offset from work_date.
  2. wk = df.groupby(['consultant_id','engagement_id','week_start'], as_index=False).hours.sum(); wk = wk[wk.hours > 0]; wk['wi'] = (wk.week_start - pd.Timestamp('1970-01-05')).dt.days // 7.
  3. wk = wk.sort_values(['consultant_id','engagement_id','wi']); d = wk.groupby(['consultant_id','engagement_id']).wi.diff(); wk['new_spell'] = d.isna() | (d > 3); wk['spell_id'] = wk.new_spell.cumsum().
  4. spells = wk.groupby(['consultant_id','engagement_id','spell_id']).agg(start_week=('week_start','min'), end_week=('week_start','max'), active_weeks=('wi','size'), total_hours=('hours','sum'), span=('wi', lambda s: s.max() - s.min() + 1)).reset_index(); gap_weeks = span - active_weeks.
  5. Run the conservation assertion (spell hours sum to input hours) and the toy case below before returning.
EXPECTED RESULTOne row per spell, where a single pair with active week indices [1, 2, 5, 9] produces two spells: weeks 1 to 5 with active_weeks 3 and gap_weeks 2, and week 9 alone with active_weeks 1 and gap_weeks 0.
Follow-up
  • Re-run with a one-week and a four-week tolerance. What happens to the spell count, and which tolerance would you defend to a staffing lead?
  • Using these spells, how would you measure how many engagements a consultant is split across in a given week, and why is that not just a count of rows?

Kaplan-Meier days to payment with unpaid invoices censored

hard
survival analysiscensoringkaplan-meier

invoice_lines has invoice_line_id, engagement_id, line_type, issued_at, due_date, paid_at (null when unpaid), amount_usd and status in draft, issued, partially_paid, paid, disputed, written_off. At a given snapshot_date, estimate the median days from issue to full payment. Implement Kaplan-Meier yourself; no lifelines or equivalent. Treat issued, partially_paid and disputed as right-censored at snapshot_date, and decide and justify what to do with written_off. Report the naive mean over paid lines alongside your estimate and state the sign of its bias.

Approach
  1. Build the duration and event table explicitly. Drop draft lines, which have no clock. For status paid, duration = (paid_at - issued_at).days with event = 1. For issued, partially_paid and disputed, duration = (snapshot_date - issued_at).days with event = 0.
  2. Handle written_off as a competing event rather than a censor. Censoring it makes the estimator answer 'time to payment if written-off invoices could still pay', which overstates collection. Either report a cumulative-incidence version alongside, or censor them and say plainly that the result is conditional on eventual collection.
  3. Implement the estimator directly: sort unique event times, at each t take n_i as the count with duration >= t and d_i as the payments at exactly t, and accumulate S(t) = product of (1 - d_i / n_i). Censored rows leave the risk set without causing a drop, which is the whole mechanism and the reason the answer differs from any completed-case average.
  4. Read the median as min{t : S(t) <= 0.5}. If S never reaches 0.5 within observed follow-up, report 'not reached'; interpolating past the last observation invents data that the snapshot does not contain.
  5. Add Greenwood's formula for Var(S(t)) to put a band on the curve, then invert the band at 0.5 for an interval on the median rather than quoting a point estimate alone.
  6. Compare against the mean over paid lines only and name the direction: at any snapshot the paid set over-represents fast payers, so the naive mean is biased low, and the bias widens exactly when collections deteriorate.
Follow-up
  • A large account moved to a monthly payment run. Is administrative censoring still independent of payment time, and what would you check?
  • Finance wants one DSO number against a target. What do you give them, and what do you refuse to give them?
  • Stratify by line_type. Do milestone lines behave like fees lines, and what would it mean for the firm if they do not?

Instead of guessing where the week should go, day one measures it under a fixed rubric and allocates the remaining hours in proportion to the gaps. The method is deliberately rigid: the allocation is written down before any studying starts and is not renegotiated when a topic turns out to be unpleasant.

Small steps. Visible outcomes.0 / 7 completed
ONE WEEK · YOUR PACE

Prepare, practise & reflect

One practical outcome each day. Spend longer where you need it.

0 / 7 done
01Diagnostic, scored before you study anything
  • Sit a 100-minute timed diagnostic in four blocks: 30 minutes of SQL across three prompts, 25 minutes of short-answer statistics, 25 minutes on one modelling or case prompt, and 20 minutes delivering one behavioural story aloud.
  • Score each block from 0 to 3 on a fixed rubric where 3 is correct and fluent, 2 is correct but slow or prompted, 1 is partially correct, and 0 is stuck, grading the output rather than how the attempt felt.
  • Allocate the hours for days two to five roughly in proportion to 3 minus the score in each block, write the allocation down, and commit to not revising it midweek.

Deliverable: A scored rubric and a fixed hour allocation for the rest of the week.

Practice prompt ↗Practice prompt ↗Worked solution ↗
02Largest gap: find the boundary rather than the subject
  • Break the weakest area into five named sub-skills (for query work: grain control, window frames, date arithmetic, set logic with NULLs, and reading a query plan) and rate each one, so the rest of the week targets a sub-skill instead of a subject.
  • Solve three problems chosen to sit just above where the rating drops off, and for each write the first move you failed to make.
  • Re-solve one of them from memory four hours later, on paper, with nothing open.

Deliverable: A five-item sub-skill map with the two blocking sub-skills circled.

Practice prompt ↗Practice prompt ↗
03Largest gap: drill the blocking sub-skill
  • Do eight short repetitions of the same shape rather than eight different problems, so what you practise is the pattern and not the puzzle.
  • Write the rule you now hold in one sentence, then test it against a case built to break it: a ranking function over a column with ties, or a two-sample test on observations that are obviously dependent.
  • Have someone else read your one-sentence rule and find the precondition you left out.

Deliverable: One rule statement with its preconditions attached and one counterexample that would have caught the incomplete version.

Practice prompt ↗Practice prompt ↗
04Second gap, plus maintenance on your strongest area
  • Run the same sub-skill map and boundary protocol on the second-largest gap, compressed into half the day.
  • Spend 25 timed minutes on your strongest area to stop it decaying, choosing the hardest problem you can still finish rather than an easy warm-up.
  • Compare how the two areas fail: whether you lose time on recall, on setup, or on arithmetic, because the fix differs for each.

Deliverable: A second sub-skill map plus a one-line diagnosis of how each area fails you.

Practice prompt ↗Practice prompt ↗Worked solution ↗
05The gap that is not a skill
  • Record yourself answering one technical and one behavioural prompt, then count two things in the playback: how many seconds before your first clarifying question, and how many sentences you started without knowing where they ended.
  • Rewrite your three most-used stock phrases into shorter versions, and practise saying "I do not know, here is how I would find out" without softening it into a guess.
  • Deliver one answer again with a hard 90-second limit to force structure before detail.

Deliverable: Two recordings with a counted improvement in time-to-first-question.

Practice prompt ↗Practice prompt ↗
06Retest under day-one conditions
  • Sit the same 100-minute diagnostic structure with new prompts of comparable difficulty and score it on the identical rubric.
  • Compare block by block, and for any block that did not move, change the method rather than adding hours: a block stuck at 1 usually means the practice was too varied, not too short.
  • Write which single block you would still lose the offer on.

Deliverable: A second scored rubric placed next to the first, with one named remaining risk.

Practice prompt ↗Practice prompt ↗
07Full loop under interview conditions
  • Run a 60-minute mock covering the two blocks that moved least, with an interviewer instructed to interrupt and change direction.
  • Write your recovery script for the moment you go blank: restate the question, state your assumption, name the first thing you would check.
  • Reduce the week to the rule statements you wrote, each with its preconditions attached, then say every one of them out loud without reading it and cut any you cannot state in a single sentence, since a rule you have to reconstruct mid-answer will not survive being interrupted.

Deliverable: A one-page card holding the recovery script and only the rules you could state from memory.

Practice prompt ↗Practice prompt ↗Worked solution ↗

Expand any day for tasks and deliverables. Your progress is saved on this device.

Most of the questions in this section reduce to one thing: can you be handed a vague request and come back with something useful? Prepare an example where the ask was underspecified, you chose an interpretation, and you said out loud which interpretation you chose. Describing how you narrowed the question matters more than the technique you eventually used.

How do you handle missing values or nulls in a production-level SQL pi…

medium
behavioural and stakeholder questions

How do you handle missing values or nulls in a production-level SQL pipeline?

Approach
  1. State the situation in two sentences and spend the rest on your reasoning.
  2. Quantify the outcome, including what you would not claim credit for.
  3. Name the disagreement or constraint, and how you resolved it with evidence.
Follow-up
  • What would you do differently if you ran that project again?
  • What did you decide not to do, and why?

Scoping a one-line question about practice profitability in twenty minutes

medium
scopingmetric definitionexclusions

A practice leader asks, in one line, whether the data_and_ai engagements are profitable. You have fct_engagement, fct_time_entry and fct_invoice_line, and twenty minutes before they travel. Some engagements are retainer or outcome_based, so contracted_hours is NULL; some invoices are unpaid; pursuit hours sit on charge_code = 'client_pursuit' with engagement_id NULL. Produce the scoping conversation: the questions you ask, the metric you commit to in writing, the exclusions you will have to make, and what you deliver by when.

Approach
  1. Name the probe: whether you convert an ambiguous one-liner into a decision with a metric attached, or start querying and return a number nobody can act on.
  2. Ask what changes based on the answer. Repricing the rate card, changing the staffing pyramid, and deciding whether to keep selling fixed_fee are three different questions, and they need different cuts of the same data. Pick the one the leader actually has authority over.
  3. Commit the metric in writing on the spot: fees plus milestone plus credit_note lines from fct_invoice_line as the revenue side, the sum of hours times cost_rate_usd over all approved entries, billable and non-billable alike, as the cost side, inception to date, reported separately by pricing_model.
  4. List what you will not answer and why: overrun on retainer and outcome_based work, because contracted_hours is NULL there and needs an explicit predicate rather than a silent NULL drop; cash collected, because unpaid invoices are censored; and pursuit cost, which has no engagement_id and must either be allocated on a stated rule or reported as a separate line.
  5. Commit two deliverables with dates rather than one vague one: the stratified margin table within two days, the mix decomposition within a week. Name the single assumption that, if wrong, flips the conclusion.
Follow-up
  • They insist on one number for a partner meeting. Which single number do you give, and what sentence goes with it?
  • If you had to allocate client_pursuit hours to engagements, what rule would you use and how would you show the answer is not sensitive to it?

Your utilisation dashboard caused a staffing decision on incomplete data

medium
late-arriving datapostmortemreporting windows

Six weeks ago you shipped a weekly billable-utilisation dashboard keyed on fct_time_entry.work_date. It showed a six-point drop across the three most recent weeks. A practice lead pulled two consultants off an engagement in response. The drop reversed on the next refresh, because time entries are created days to weeks after the work happens and entered_at trails work_date. Describe what you do now: the diagnosis, the change to the artifact, and the conversation with the person who acted on your number.

Approach
  1. Name the probe: whether you own a reporting-design error rather than reclassifying it as someone else's timesheet compliance problem, and whether you fix the class of bug instead of the single week.
  2. Quantify before explaining. Measure the backfill curve directly from entered_at: for work_date D, the share of final hours that existed as of D plus k, for k from 1 to 45. That gives an observed tail length instead of a guessed cutoff.
  3. Change the artifact so the incomplete region cannot be read as a trend. Either end the trended series at snapshot minus the measured tail, or publish an as-of-entered_at series that is internally consistent, and label which one is on screen.
  4. Tell the person who acted, first and directly, with the corrected series and the specific decision to revisit. A correction that arrives after they notice costs more than the original error.
  5. Add a standing completeness tile: timesheet submission lag, the share of entries where entered_at date minus work_date exceeds seven days, so the dashboard shows its own reliability rather than depending on you remembering.
Follow-up
  • Leadership still wants to see the current week. What do you show, and how do you label it?
  • One practice runs a three-day lag and another twenty days. Do you set one firm-wide cutoff or one per practice, and what does that cost in comparability?
  • 01

    How do you handle missing values or nulls in a production-level SQL pipeline?

  • 02

    A practice leader asks, in one line, whether the data_and_ai engagements are profitable. You have fct_engagement, fct_time_entry and fct_invoice_line, and twenty minutes before they travel. Some engagements are retainer or outcome_based, so contracted_hours is NULL; some invoices are unpaid; pursuit hours sit on charge_code = 'client_pursuit' with engagement_id NULL. Produce the scoping conversation: the questions you ask, the metric you commit to in writing, the exclusions you will have to make, and what you deliver by when.

  • 03

    Six weeks ago you shipped a weekly billable-utilisation dashboard keyed on fct_time_entry.work_date. It showed a six-point drop across the three most recent weeks. A practice lead pulled two consultants off an engagement in response. The drop reversed on the next refresh, because time entries are created days to weeks after the work happens and entered_at trails work_date. Describe what you do now: the diagnosis, the change to the artifact, and the conversation with the person who acted on your number.

PracHub interview preparation framework ↗
Is this an official Tata Consultancy Services interview guide?

No. It is PracHub's own research and practice material for the Data Scientist role at Tata Consultancy Services. Rounds and questions reflect what candidates have reported, not a process Tata Consultancy Services has published, and they change over time. Confirm the current format and scope with your recruiter.

PracHub interview research ↗
How long should I prepare for the technical rounds?

Most candidates spend 4–6 weeks of consistent practice. Focus on mastering core SQL functions and reviewing statistical fundamentals rather than memorizing specific solutions.

PracHub interview research ↗
What is the most common reason candidates fail the technical screen?

Lack of clarity in communication. Even if your code is correct, you must explain your logic and why you chose a specific approach over alternatives.

PracHub interview research ↗
Is the interview process mostly remote or in-person?

Tata Consultancy Services often utilizes a mix of both. Be prepared for virtual coding platforms and video conferencing, but remain flexible regarding local office requirements.

PracHub interview research ↗
How much weight is given to behavioral questions?

Behavioral rounds are critical. They determine whether you can work effectively within a team and handle the pressure of client-facing projects.

PracHub interview research ↗
Sources & methodology 3 sources ↗

Official role evidence, timestamped platform data and clearly labeled preparation advice.