A Data Scientist at Tata Consultancy Services (TCS) functions at the intersection of advanced analytics, business strategy, and technical implementation. You will be responsible for translating complex business requirements into actionable data models, providing insights that drive decision-making for large-scale enterprise clients. Your work directly impacts how organizations optimize their operations, enhance customer experiences, and leverage predictive modeling to maintain a competitive edge.
This role is critical to the Tata Consultancy Services mission of providing high-value IT solutions. You will engage in the full lifecycle of data science projects, from raw data extraction and exploratory analysis to the deployment of machine learning models. Whether you are working on supply chain optimization, churn prediction, or personalized customer analytics, you will be expected to balance technical rigor with the practical realities of product-driven business environments.
Preparation focus
editorialNo round sequence has been reported for this company, so work the categories below and confirm the format with your recruiter.
What to demonstrate
- Breadth across SQL, experimentation and product reasoning
- Ability to state assumptions before choosing a method
How to prepare
- Drill the practice exercises below and time yourself
- Prepare three quantified stories about decisions you drove
PracHub editorial advice for the preparation topics above.
Pooling margin, realisation or overrun across pricing models
Fixed-fee margin falls with hours worked; uncapped time-and-materials margin rises with hours worked; retainer margin depends on neither. A quarter in which the firm sells more fixed-fee work will show a margin change caused entirely by mix, not by delivery performance, and the aggregate can move in the opposite direction to every individual pricing model. Always stratify by fct_engagement.pricing_model before comparing periods, and report the mix shift alongside the within-stratum change.
Trending utilisation or revenue on work_date without accounting for timesheet backfill
Time entries are created days to weeks after the work happens, and the backfill tail often runs two to six weeks. A dashboard keyed on work_date therefore shows the most recent weeks as a decline that reverses on every refresh. The fix is either to hold the reporting window back past the observed backfill tail (measure the tail with the timesheet submission lag metric rather than guessing) or to report an as-of-entered_at snapshot so the series is internally consistent, and to state which one you used.
Over-explaining the method and under-explaining the implication
Lead with the answer and what you would do about it, then give the approach when asked. Roughly one sentence of method per three of implication is the right ratio for a stakeholder-facing answer; the interviewer already knows what a regression is.
Reporting a p-value with no effect size or interval
Give the estimated difference with a confidence interval in the units the business cares about, then say whether that whole interval is worth acting on. A p-value only addresses whether you can rule out exactly zero; it says nothing about magnitude.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Cluster bootstrap a margin change across few, unequal accounts
engagements has engagement_id, client_id, pricing_model, fees_usd, cost_usd and period in {pre, post} around a rate-card change. There are about 35 client_ids across roughly 400 engagements, and the top five clients carry a large share of fees. Estimate the pre-to-post change in value-weighted gross margin within pricing_model, and give a 95% interval by resampling whole client_ids with replacement, B = 2000. Write the bootstrap yourself; no resampling library. Report the account-weighted estimate alongside and explain any divergence.
Approach
- Write the statistic as a pure function of a dataframe first. Within each pricing_model, margin = (sum(fees) - sum(cost)) / sum(fees) per period; the headline is the fee-weighted average of the within-model differences using fixed pre-period weights, so a shift in the mix of work sold cannot masquerade as a change in delivery.
- Resample clusters, not rows: draw 35 client_ids with replacement and concatenate all engagements of each drawn client. A client drawn twice contributes its rows twice under distinct pseudo-ids, which is what preserves within-account correlation instead of averaging it away.
- Recompute the statistic on each replicate and take a percentile interval from the 2.5th and 97.5th quantiles. With concentrated fees the replicate distribution is skewed, so a symmetric point plus or minus 1.96 times a standard error is wrong in the tail that matters.
- Compute the account-weighted version (mean across clients of each client's margin change) next to the value-weighted one, and attribute the divergence to concentration rather than to noise; if they disagree in sign, that fact is the finding.
- Name the regime honestly: at roughly 35 clusters both cluster-robust standard errors and the pairs cluster bootstrap under-cover, so state the fix you would run next, which is a wild cluster bootstrap with Rademacher weights or CR2 with t(G-1) critical values.
Follow-up
- Implement the wild cluster bootstrap and show at what cluster count its p-value separates from the naive one.
- One client is 30% of fees. Show the interval with and without that account and say which one you would present, and to whom.
- What pre-period check would make you willing to call this a causal effect of the rate card rather than a correlation?
Collapse time entries into contiguous staffing spells
From approved delivery time entries (consultant_id, engagement_id, work_date, hours, charge_code, status, with engagement_id not null), build staffing spells. Per consultant and engagement, collapse weeks containing any logged hours into contiguous runs, where three or more consecutive zero-hour weeks end a spell. Output one row per spell with consultant_id, engagement_id, start_week, end_week, active_weeks, gap_weeks and total_hours. Consultants sit on several engagements at once, so spells from different engagements may overlap in time and must not be merged. No Python loop over rows.
Approach
- Aggregate to (consultant_id, engagement_id, week_start) with summed hours and keep only weeks with positive hours. The absent weeks are the signal, so materialising zeros here would destroy the thing you are detecting.
- Convert week_start into an integer week index, ((week_start - epoch_monday).dt.days // 7), so gap detection is integer subtraction rather than calendar arithmetic that breaks over month and year boundaries.
- Sort by [consultant_id, engagement_id, week_index], diff the week index inside each pair, and mark a spell start where the diff is null (first row of the pair) or greater than 3. A diff of 1 is adjacent weeks and a diff of 3 is two empty weeks, which the tolerance permits.
- Take spell_id = the cumulative sum of that boolean over the whole frame so ids are globally unique, then a single groupby on [consultant_id, engagement_id, spell_id] yields min and max week, active week count and summed hours; gap_weeks = (end - start + 1) - active_weeks.
- Do not deduplicate overlapping spells across engagements. A consultant on two engagements in the same week is the normal case, and that overlap is the fact any capacity or context-switching question needs.
Worked solution 30 min
- Filter to status == 'approved', charge_code == 'client_delivery' and engagement_id.notna(); derive week_start by subtracting the weekday offset from work_date.
- wk = df.groupby(['consultant_id','engagement_id','week_start'], as_index=False).hours.sum(); wk = wk[wk.hours > 0]; wk['wi'] = (wk.week_start - pd.Timestamp('1970-01-05')).dt.days // 7.
- wk = wk.sort_values(['consultant_id','engagement_id','wi']); d = wk.groupby(['consultant_id','engagement_id']).wi.diff(); wk['new_spell'] = d.isna() | (d > 3); wk['spell_id'] = wk.new_spell.cumsum().
- spells = wk.groupby(['consultant_id','engagement_id','spell_id']).agg(start_week=('week_start','min'), end_week=('week_start','max'), active_weeks=('wi','size'), total_hours=('hours','sum'), span=('wi', lambda s: s.max() - s.min() + 1)).reset_index(); gap_weeks = span - active_weeks.
- Run the conservation assertion (spell hours sum to input hours) and the toy case below before returning.
Follow-up
- Re-run with a one-week and a four-week tolerance. What happens to the spell count, and which tolerance would you defend to a staffing lead?
- Using these spells, how would you measure how many engagements a consultant is split across in a given week, and why is that not just a count of rows?
Kaplan-Meier days to payment with unpaid invoices censored
invoice_lines has invoice_line_id, engagement_id, line_type, issued_at, due_date, paid_at (null when unpaid), amount_usd and status in draft, issued, partially_paid, paid, disputed, written_off. At a given snapshot_date, estimate the median days from issue to full payment. Implement Kaplan-Meier yourself; no lifelines or equivalent. Treat issued, partially_paid and disputed as right-censored at snapshot_date, and decide and justify what to do with written_off. Report the naive mean over paid lines alongside your estimate and state the sign of its bias.
Approach
- Build the duration and event table explicitly. Drop draft lines, which have no clock. For status paid, duration = (paid_at - issued_at).days with event = 1. For issued, partially_paid and disputed, duration = (snapshot_date - issued_at).days with event = 0.
- Handle written_off as a competing event rather than a censor. Censoring it makes the estimator answer 'time to payment if written-off invoices could still pay', which overstates collection. Either report a cumulative-incidence version alongside, or censor them and say plainly that the result is conditional on eventual collection.
- Implement the estimator directly: sort unique event times, at each t take n_i as the count with duration >= t and d_i as the payments at exactly t, and accumulate S(t) = product of (1 - d_i / n_i). Censored rows leave the risk set without causing a drop, which is the whole mechanism and the reason the answer differs from any completed-case average.
- Read the median as min{t : S(t) <= 0.5}. If S never reaches 0.5 within observed follow-up, report 'not reached'; interpolating past the last observation invents data that the snapshot does not contain.
- Add Greenwood's formula for Var(S(t)) to put a band on the curve, then invert the band at 0.5 for an interval on the median rather than quoting a point estimate alone.
- Compare against the mean over paid lines only and name the direction: at any snapshot the paid set over-represents fast payers, so the naive mean is biased low, and the bias widens exactly when collections deteriorate.
Follow-up
- A large account moved to a monthly payment run. Is administrative censoring still independent of payment time, and what would you check?
- Finance wants one DSO number against a target. What do you give them, and what do you refuse to give them?
- Stratify by line_type. Do milestone lines behave like fees lines, and what would it mean for the firm if they do not?
Describe your process for optimizing a slow-running query on a massive…
Describe your process for optimizing a slow-running query on a massive dataset.
Approach
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- Say which table is the grain you start from, and join outward from it.
- Check whether any join is one-to-many before aggregating, or the sums inflate.
Follow-up
- How does the query change if the join becomes one-to-many?
- What breaks if events arrive late or out of order?
How would you use SQL window functions to calculate a running total or…
How would you use SQL window functions to calculate a running total or a moving average?
Approach
- State the window function and its partition and ordering out loud before writing it.
- Check whether any join is one-to-many before aggregating, or the sums inflate.
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
Follow-up
- How would you verify this result without re-running the same query?
- What breaks if events arrive late or out of order?
Write a query to identify the top three customers by spend per region.
Write a query to identify the top three customers by spend per region.
Approach
- State the window function and its partition and ordering out loud before writing it.
- Check whether any join is one-to-many before aggregating, or the sums inflate.
- Say which table is the grain you start from, and join outward from it.
Follow-up
- How does the query change if the join becomes one-to-many?
- What breaks if events arrive late or out of order?
Explain the performance trade-offs between using a join versus a subqu…
Explain the performance trade-offs between using a join versus a subquery in large databases.
Approach
- State the window function and its partition and ordering out loud before writing it.
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- Check whether any join is one-to-many before aggregating, or the sums inflate.
Follow-up
- What breaks if events arrive late or out of order?
- How would you verify this result without re-running the same query?
Find signed engagements with no approved billable hour
fct_engagement holds engagement_id, start_date and status. fct_time_entry holds time_entry_id, consultant_id, engagement_id, charge_code, work_date, hours, is_billable and status, and engagement_id is NULL on bench, leave, training and internal rows. Return engagements with start_date in the trailing quarter and status IN ('active','closed_delivered','closed_early') that have no approved billable time entry with work_date within 45 days of start_date, giving engagement_id, start_date and days elapsed. Then explain why NOT IN over a subquery selecting fct_time_entry.engagement_id is the wrong instrument here.
Approach
- Express the exclusion as NOT EXISTS with a correlated predicate covering all four conditions at once: te.engagement_id = e.engagement_id AND te.is_billable AND te.status = 'approved' AND te.work_date < e.start_date + INTERVAL '45 days'. NOT IN can only carry a single column, so a multi-condition anti-join forces a contorted subquery even before the NULL problem.
- State the NULL semantics precisely: NOT IN (SELECT engagement_id FROM fct_time_entry) evaluates to NULL rather than TRUE for every candidate as soon as one NULL is present in the list, because x <> NULL is unknown. The WHERE clause keeps only TRUE, so the query returns zero rows and reads as a clean bill of health. This table guarantees NULLs, since bench and leave rows carry no engagement.
- Note the equivalent LEFT JOIN ... WHERE te.engagement_id IS NULL form. It is correct, but it materialises a join only to discard most of it; NOT EXISTS states the intent and most planners execute both as the same anti-join.
- Guard the window against truncation: an engagement that started 10 days ago cannot yet have failed a 45-day test, so require start_date <= current_date - 45 or report the too-early set separately rather than counting them as failures.
- Decide on the late-timesheet risk explicitly. A missing entry may mean unstaffed work or an unsubmitted timesheet, so either hold the cutoff back past the measured submission lag or label the recent tail as provisional.
Worked solution 20 min
- Count engagements matching the status and start_date predicates; this is the denominator.
- Write the NOT EXISTS version and record its row count.
- Write the NOT IN version on the same data and observe zero rows, then confirm with SELECT COUNT(*) FROM fct_time_entry WHERE engagement_id IS NULL that NULLs are present.
- Cross-check by counting engagements that DO have a qualifying entry and verifying the two counts sum to the denominator.
- Open one engagement from the result and look at its raw time entries to see whether the cause is no staffing or unapproved timesheets.
Follow-up
- Turn this into a time-to-first-billable-hour distribution. What do you do with engagements that never recorded one?
- Should client_pursuit hours count as the engagement being staffed, and how does that choice change the list?
- How would you prove the empty result from the NOT IN version is a bug rather than good news?
Diagnose a sample ratio mismatch before reading the result
A staffing-recommendation tool was randomised one to one across 840 billable-role consultants, with assignment drawn from a hash of consultant_id, so 420 sit in each arm and the assigned arm is recoverable for every consultant. The analysis table, built from exposure logs, holds 300 treatment and 420 control consultants: 720 rows against 840 assignments. Treatment shows a 1.9-point utilisation lift at p = 0.01, and the sponsor wants it in the deck tomorrow. Run the right test on the allocation, quantify how far the split sits from 1:1, decide whether the lift can be reported, and name the two most likely mechanisms given that the table came from exposure logs rather than the assignment table.
Approach
- Test the split before touching the outcome. Chi-square goodness of fit on the 720 retained rows against an expected 360/360 gives chi-square = 2 * (60^2 / 360) = 20.0 on 1 degree of freedom, p around 7.8e-6. The equivalent binomial z is 60 / sqrt(720 * 0.25) = 4.47, and z^2 = 20.0 reproduces the chi-square.
- Then run the sharper test that a recoverable assignment makes available: per-arm retention against the known 420. Control retains 420/420 = 100%, treatment retains 300/420 = 71.4%. That 28.6-point gap says more than the split test does, because it names which arm lost rows and how many, rather than only that the ratio is off.
- Use a much stricter alpha for the allocation test than for the outcome, commonly 0.0005. The split test protects no multiplicity budget and should fire only on real breakage; 7.8e-6 clears even that bar by a factor of about sixty.
- Treat the outcome as unreportable rather than as interesting-with-a-caveat. Under a sample ratio mismatch the analysis population is no longer the randomised population, so the comparison is observational and the lift carries no protection from confounding.
- Trace the mechanism from the table's provenance, starting from the fact that the deficit is one-sided. A symmetric filter, such as an inner join to dim_consultant on is_current = TRUE, drops rows from both arms at similar rates: it shrinks n without moving the ratio, so it cannot be the cause here and should be ruled out rather than listed. Two one-sided candidates fit. First, the exposure log records a treatment consultant only once they open the tool, so non-openers vanish from treatment and the arm is silently conditioned on engagement, which is a post-assignment filter. Second, the arms are populated by different loggers: treatment exposure is written by the tool's own client-side beacon while control exposure is written server-side by the staffing screen, so the arms differ in logging coverage rather than in behaviour.
- Rebuild on intention to treat from the assignment table, so all 840 assigned consultants sit in their assigned arm whether or not they opened the tool, then re-run the split test before re-running the outcome.
Follow-up
- Suppose the analysis table had held all 840 consultants, split 402 to 438. Run the same test and tell me whether you would report the result.
- If non-openers are themselves a treatment effect, because the tool is bad enough that people ignore it, what does intention to treat estimate, and what would you need to estimate the effect on compliers?
- How would you monitor for sample ratio mismatch continuously without creating a peeking problem on the outcome?
Estimate a mandatory pricing review at a value threshold
Proposals with expected_value_usd at or above 250,000 must pass a mandatory pricing review before submission; below that the originating partner prices unilaterally. The rule has stood for three years, nothing was randomised, and leadership wants to know whether the review improves realisation on the resulting engagements. About 900 proposals a year land in fct_proposal, and realisation is only observable for those that reach stage = won. Design the identification strategy, state the estimand precisely, and name the two checks that would make you abandon the design.
Approach
- Set up the design. Running variable is expected_value_usd as recorded when the rule is applied, cutoff 250,000, outcome realisation on the resulting engagement. Estimate with local linear regression either side, triangular kernel, MSE-optimal bandwidth and robust bias-corrected confidence intervals. Do not fit a global high-order polynomial, which drags weight from proposals far from the cutoff onto the estimate at it.
- Check compliance before calling it sharp. If some sub-threshold proposals are reviewed voluntarily or some above-threshold ones skip review, the design is fuzzy, which is instrumental variables with the threshold indicator as the instrument. The estimate is a Wald ratio, the jump in realisation divided by the jump in review probability, and it identifies the effect for compliers at the cutoff only.
- Run the manipulation test and interpret it carefully. A density test on expected_value_usd will very likely reject, because proposal values heap at round numbers and 250,000 is exactly such a number. Heaping alone is not sorting, but a partner pricing at 249,000 to dodge review is, and both look the same in the density. Separate them by checking whether excess mass appears only just below the cutoff or at round numbers throughout the range, and by running a donut specification that drops a window around the threshold.
- Check covariate continuity at the cutoff for industry_segment, account_tier, is_competitive, practice_area and originating_partner_id, using the same local linear specification with each covariate as the outcome. A jump in any of them at 250,000 is a jump in the population rather than in the treatment, and it kills the design.
- Confront the outcome-side selection. Realisation exists only for won proposals, and review may itself move the win rate, so conditioning on a win conditions on a post-treatment variable. Estimate the RD on win rate first; if it jumps, switch to an outcome defined over all proposals, such as realised fees per proposal with losses entering as zero, or report Lee bounds rather than a point estimate.
- Size the design before committing to it. Only proposals inside the bandwidth contribute, so 2,700 proposals over three years with a bandwidth capturing perhaps 15% leaves an effective sample near 400 split across the cutoff. Compute the MDE on that number, not on 2,700, and say in advance whether it can detect the effect that would change the policy.
Worked solution 45 min
- Histogram expected_value_usd in fine bins around 250,000 and across the full range, to separate generic round-number heaping from cutoff-specific sorting.
- Estimate the first stage: probability of a recorded review as a function of the running variable, and read the jump at the cutoff. A jump well below one means fuzzy, so move to the Wald ratio.
- Run the local linear RD on realisation with an MSE-optimal bandwidth and robust bias-corrected intervals, then repeat at half and twice that bandwidth.
- Repeat the same specification with each pre-treatment covariate as the outcome, and with placebo cutoffs at 200,000 and 300,000.
- Run the RD on win rate, then on realised fees per proposal defined over all proposals, and compare the three results.
- Report the MDE computed on the in-bandwidth proposal count alongside the estimate.
Follow-up
- If the density test rejects and the donut specification moves the estimate materially, what do you conclude and what do you put in the deck?
- The estimand is local to 250,000 and leadership wants to know whether to lower the threshold to 100,000. What can you honestly say?
- The review is triggered on the value at submission, but the value is revised afterwards. Which version of the running variable do you use, and why?
Days sales outstanding improved while collections got worse
A cash dashboard reports mean days-to-pay as AVG(paid_at - issued_at) over fct_invoice_line where status = 'paid'. It improved from 47 to 39 days last quarter. The finance lead insists collections are worse, and invoice volume rose about 30% after a large batch issued early in the quarter. Using invoice_line_id, engagement_id, issued_at, due_date, paid_at, status and amount_usd, explain the contradiction and deliver the estimate you would report instead, with its assumptions.
Approach
- Name the bias precisely: at any snapshot, the set of paid invoices over-represents fast payers, because slow ones have not finished yet. Unpaid, partially paid and disputed lines are right-censored observations, not absent ones, and averaging over completions alone is biased downward.
- Explain why the bias grew: a large young cohort adds many invoices whose slow half cannot yet appear in the paid set, so a volume increase alone pushes the naive mean down even with unchanged payment behaviour.
- Build the censored dataset: event time = paid_at - issued_at for paid lines; censoring time = snapshot_date - issued_at for status IN ('issued','partially_paid','disputed'). Confirm no placeholder dates were substituted for NULL paid_at.
- Estimate with Kaplan-Meier by issue-month cohort and report both the median and the restricted mean to a fixed horizon such as 90 days, so cohorts of different maturity are compared over the same span. State the independent-censoring assumption: censoring here is administrative, driven by issue date rather than by payment behaviour, which is what makes it defensible.
- Cross-check with a metric that needs no model: share of each cohort's invoiced amount collected by day 30, 60 and 90. If the recent cohort's curve sits below prior cohorts at the same age, collections really did deteriorate.
- Report amount-weighted as well as count-weighted, since one large disputed invoice moves cash without moving the count.
Follow-up
- Disputed invoices may never pay at all. Does that break the censoring assumption, and how would you handle it?
- How would you turn this into a weekly operational report that cannot be gamed by issuing more invoices?
- Proposal cycle time has the same structure. What is the equivalent fix there?
Instead of guessing where the week should go, day one measures it under a fixed rubric and allocates the remaining hours in proportion to the gaps. The method is deliberately rigid: the allocation is written down before any studying starts and is not renegotiated when a topic turns out to be unpleasant.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Diagnostic, scored before you study anything
- Sit a 100-minute timed diagnostic in four blocks: 30 minutes of SQL across three prompts, 25 minutes of short-answer statistics, 25 minutes on one modelling or case prompt, and 20 minutes delivering one behavioural story aloud.
- Score each block from 0 to 3 on a fixed rubric where 3 is correct and fluent, 2 is correct but slow or prompted, 1 is partially correct, and 0 is stuck, grading the output rather than how the attempt felt.
- Allocate the hours for days two to five roughly in proportion to 3 minus the score in each block, write the allocation down, and commit to not revising it midweek.
Deliverable: A scored rubric and a fixed hour allocation for the rest of the week.
Practice prompt ↗Practice prompt ↗Worked solution ↗02Largest gap: find the boundary rather than the subject
- Break the weakest area into five named sub-skills (for query work: grain control, window frames, date arithmetic, set logic with NULLs, and reading a query plan) and rate each one, so the rest of the week targets a sub-skill instead of a subject.
- Solve three problems chosen to sit just above where the rating drops off, and for each write the first move you failed to make.
- Re-solve one of them from memory four hours later, on paper, with nothing open.
Deliverable: A five-item sub-skill map with the two blocking sub-skills circled.
Practice prompt ↗Practice prompt ↗03Largest gap: drill the blocking sub-skill
- Do eight short repetitions of the same shape rather than eight different problems, so what you practise is the pattern and not the puzzle.
- Write the rule you now hold in one sentence, then test it against a case built to break it: a ranking function over a column with ties, or a two-sample test on observations that are obviously dependent.
- Have someone else read your one-sentence rule and find the precondition you left out.
Deliverable: One rule statement with its preconditions attached and one counterexample that would have caught the incomplete version.
Practice prompt ↗Practice prompt ↗04Second gap, plus maintenance on your strongest area
- Run the same sub-skill map and boundary protocol on the second-largest gap, compressed into half the day.
- Spend 25 timed minutes on your strongest area to stop it decaying, choosing the hardest problem you can still finish rather than an easy warm-up.
- Compare how the two areas fail: whether you lose time on recall, on setup, or on arithmetic, because the fix differs for each.
Deliverable: A second sub-skill map plus a one-line diagnosis of how each area fails you.
Practice prompt ↗Practice prompt ↗Worked solution ↗05The gap that is not a skill
- Record yourself answering one technical and one behavioural prompt, then count two things in the playback: how many seconds before your first clarifying question, and how many sentences you started without knowing where they ended.
- Rewrite your three most-used stock phrases into shorter versions, and practise saying "I do not know, here is how I would find out" without softening it into a guess.
- Deliver one answer again with a hard 90-second limit to force structure before detail.
Deliverable: Two recordings with a counted improvement in time-to-first-question.
Practice prompt ↗Practice prompt ↗06Retest under day-one conditions
- Sit the same 100-minute diagnostic structure with new prompts of comparable difficulty and score it on the identical rubric.
- Compare block by block, and for any block that did not move, change the method rather than adding hours: a block stuck at 1 usually means the practice was too varied, not too short.
- Write which single block you would still lose the offer on.
Deliverable: A second scored rubric placed next to the first, with one named remaining risk.
Practice prompt ↗Practice prompt ↗07Full loop under interview conditions
- Run a 60-minute mock covering the two blocks that moved least, with an interviewer instructed to interrupt and change direction.
- Write your recovery script for the moment you go blank: restate the question, state your assumption, name the first thing you would check.
- Reduce the week to the rule statements you wrote, each with its preconditions attached, then say every one of them out loud without reading it and cut any you cannot state in a single sentence, since a rule you have to reconstruct mid-answer will not survive being interrupted.
Deliverable: A one-page card holding the recovery script and only the rules you could state from memory.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Most of the questions in this section reduce to one thing: can you be handed a vague request and come back with something useful? Prepare an example where the ask was underspecified, you chose an interpretation, and you said out loud which interpretation you chose. Describing how you narrowed the question matters more than the technique you eventually used.
How do you handle missing values or nulls in a production-level SQL pi…
How do you handle missing values or nulls in a production-level SQL pipeline?
Approach
- State the situation in two sentences and spend the rest on your reasoning.
- Quantify the outcome, including what you would not claim credit for.
- Name the disagreement or constraint, and how you resolved it with evidence.
Follow-up
- What would you do differently if you ran that project again?
- What did you decide not to do, and why?
Scoping a one-line question about practice profitability in twenty minutes
A practice leader asks, in one line, whether the data_and_ai engagements are profitable. You have fct_engagement, fct_time_entry and fct_invoice_line, and twenty minutes before they travel. Some engagements are retainer or outcome_based, so contracted_hours is NULL; some invoices are unpaid; pursuit hours sit on charge_code = 'client_pursuit' with engagement_id NULL. Produce the scoping conversation: the questions you ask, the metric you commit to in writing, the exclusions you will have to make, and what you deliver by when.
Approach
- Name the probe: whether you convert an ambiguous one-liner into a decision with a metric attached, or start querying and return a number nobody can act on.
- Ask what changes based on the answer. Repricing the rate card, changing the staffing pyramid, and deciding whether to keep selling fixed_fee are three different questions, and they need different cuts of the same data. Pick the one the leader actually has authority over.
- Commit the metric in writing on the spot: fees plus milestone plus credit_note lines from fct_invoice_line as the revenue side, the sum of hours times cost_rate_usd over all approved entries, billable and non-billable alike, as the cost side, inception to date, reported separately by pricing_model.
- List what you will not answer and why: overrun on retainer and outcome_based work, because contracted_hours is NULL there and needs an explicit predicate rather than a silent NULL drop; cash collected, because unpaid invoices are censored; and pursuit cost, which has no engagement_id and must either be allocated on a stated rule or reported as a separate line.
- Commit two deliverables with dates rather than one vague one: the stratified margin table within two days, the mix decomposition within a week. Name the single assumption that, if wrong, flips the conclusion.
Follow-up
- They insist on one number for a partner meeting. Which single number do you give, and what sentence goes with it?
- If you had to allocate client_pursuit hours to engagements, what rule would you use and how would you show the answer is not sensitive to it?
Your utilisation dashboard caused a staffing decision on incomplete data
Six weeks ago you shipped a weekly billable-utilisation dashboard keyed on fct_time_entry.work_date. It showed a six-point drop across the three most recent weeks. A practice lead pulled two consultants off an engagement in response. The drop reversed on the next refresh, because time entries are created days to weeks after the work happens and entered_at trails work_date. Describe what you do now: the diagnosis, the change to the artifact, and the conversation with the person who acted on your number.
Approach
- Name the probe: whether you own a reporting-design error rather than reclassifying it as someone else's timesheet compliance problem, and whether you fix the class of bug instead of the single week.
- Quantify before explaining. Measure the backfill curve directly from entered_at: for work_date D, the share of final hours that existed as of D plus k, for k from 1 to 45. That gives an observed tail length instead of a guessed cutoff.
- Change the artifact so the incomplete region cannot be read as a trend. Either end the trended series at snapshot minus the measured tail, or publish an as-of-entered_at series that is internally consistent, and label which one is on screen.
- Tell the person who acted, first and directly, with the corrected series and the specific decision to revisit. A correction that arrives after they notice costs more than the original error.
- Add a standing completeness tile: timesheet submission lag, the share of entries where entered_at date minus work_date exceeds seven days, so the dashboard shows its own reliability rather than depending on you remembering.
Follow-up
- Leadership still wants to see the current week. What do you show, and how do you label it?
- One practice runs a three-day lag and another twenty days. Do you set one firm-wide cutoff or one per practice, and what does that cost in comparability?
- 01
How do you handle missing values or nulls in a production-level SQL pipeline?
- 02
A practice leader asks, in one line, whether the data_and_ai engagements are profitable. You have fct_engagement, fct_time_entry and fct_invoice_line, and twenty minutes before they travel. Some engagements are retainer or outcome_based, so contracted_hours is NULL; some invoices are unpaid; pursuit hours sit on charge_code = 'client_pursuit' with engagement_id NULL. Produce the scoping conversation: the questions you ask, the metric you commit to in writing, the exclusions you will have to make, and what you deliver by when.
- 03
Six weeks ago you shipped a weekly billable-utilisation dashboard keyed on fct_time_entry.work_date. It showed a six-point drop across the three most recent weeks. A practice lead pulled two consultants off an engagement in response. The drop reversed on the next refresh, because time entries are created days to weeks after the work happens and entered_at trails work_date. Describe what you do now: the diagnosis, the change to the artifact, and the conversation with the person who acted on your number.
Is this an official Tata Consultancy Services interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at Tata Consultancy Services. Rounds and questions reflect what candidates have reported, not a process Tata Consultancy Services has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How long should I prepare for the technical rounds?
Most candidates spend 4–6 weeks of consistent practice. Focus on mastering core SQL functions and reviewing statistical fundamentals rather than memorizing specific solutions.
PracHub interview research ↗What is the most common reason candidates fail the technical screen?
Lack of clarity in communication. Even if your code is correct, you must explain your logic and why you chose a specific approach over alternatives.
PracHub interview research ↗Is the interview process mostly remote or in-person?
Tata Consultancy Services often utilizes a mix of both. Be prepared for virtual coding platforms and video conferencing, but remain flexible regarding local office requirements.
PracHub interview research ↗How much weight is given to behavioral questions?
Behavioral rounds are critical. They determine whether you can work effectively within a team and handle the pressure of client-facing projects.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22