Cohere · Software Engineer
Updated · 2026-09-24

Cohere Software Engineer
Interview Guide

THE 60-SECOND BRIEF

The source notes behind this guide place the Software Engineer role at Cohere in enterprise AI: platform work, developer products and data integrations built around large language models, with problems that span distributed systems, core platform engineering and application performance. The reported questions follow the same themes: tokenizing streaming text, managing context windows for language models, speeding up vector similarity search, serving high-volume generative AI APIs and pulling external data sources into a core platform.

This guide covers the six stages candidates report for the Cohere Software Engineer role: recruiter screening, a technical assessment, a core interview loop that includes a hiring manager conversation, live coding, system design, and a distributed systems debugging exercise. For each stage it explains what to expect and how to prepare. It also groups the reported questions by category (coding, system design, debugging, behavioral) and adds original SQL, coding and API-design drills with worked solutions.

Cohere candidates report 6 rounds · ≈ 4-6 weeks. The stages below are what candidates describe, not a published process.

Evolve APIs without breaking pinned SDK clientsKeep money in integer minor unitsBuild at-least-once pipelines with explicit deduplication horizons

39 min read

Practice 15 Software Engineer prompts
3Company bank questionsSnapshot · Sep 29, 2026 PT
2Candidate experiences ↗Read their reports
15Practice promptsAcross five skill areas
3With worked solutionsIncluded in the practice prompts

The source notes for this guide describe the Software Engineer role at Cohere as work on AI systems, developer products and enterprise platforms, much of it involving language models, automation tools and large-scale data integrations. The work covers distributed systems, core platform engineering and high-performance application code, and the notes mention close collaboration with product managers, researchers and other engineering groups.

The reported questions reflect that mix. The reported coding questions cover five tasks: tokenizing a text stream under tight memory limits, a custom data structure for managing and retrieving context windows for language models, optimizing an existing algorithm that handles high-throughput vector similarity search, a graph traversal for dependency resolution in distributed pipelines, and writing clean, production-ready Python or TypeScript for a new data transformation. The design questions deal with serving generative AI APIs at volume, rate limiting and caching, high availability during traffic spikes, monitoring across regions, and ingesting data from external sources.

Candidates describe a loop that runs from a recruiter screen and a practical technical assessment through a core loop with a hiring manager conversation, live coding and system design, and ends with a distributed systems debugging exercise. In that exercise, interviewers add new findings while you work through a simulated production issue. Prepare for that round separately from system design. Designing a system and diagnosing a live failure use different habits, and this guide gives the debugging round its own day of practice.

01

Recruiter Screening

reported

The source notes describe this as a conversation about your background, your interests and how well they align with the role. Use it to learn the format of every later stage, because the loop covers very different formats: a practical assessment, live coding, system design and a debugging simulation. Ask whether the technical assessment is a take-home or a live session, which languages are accepted, whether AI coding tools are allowed at any stage, and which team or product area the opening sits in. Then describe your background in terms of the areas the reported questions cover: distributed systems, platform work, developer-facing APIs and data integrations.

What to demonstrate

  • How clearly your background maps onto the areas the role spans: distributed systems, platform engineering, developer products and data integrations
  • Whether your stated interests and reasons for applying are specific to this role rather than generic
  • Whether you describe your own scope accurately, including the parts of a posting you have not done

How to prepare

  • Write a short spoken summary of your last two roles that names one distributed-systems or integration problem you owned and what you changed
  • Prepare a short list of format questions for the recruiter: assessment type, accepted languages, AI tool rules, and what the debugging round involves
  • Mark each line of the job posting as done, adjacent or new, and have an honest sentence ready for each adjacent line
PracHub interview research ↗
02

Technical Assessment

reported

Candidates report a practical technical assessment or a take-home assignment meant to test real engineering skill. One reported coding question asks for clean, production-ready Python or TypeScript that solves a new data transformation problem, and that is a useful guide to the standard to aim for. Hand in something a reviewer could merge: clear names, validated input, explicit handling of malformed or empty data, tests that cover the edge cases, and a short README with your assumptions and the trade-offs you made. If you are allowed to use AI tools, you still own every line and should be able to explain it in a follow-up.

What to demonstrate

  • Whether the code runs correctly on inputs beyond the examples provided, including empty, malformed and duplicate records
  • Whether the structure and tests would hold up in a code review, not just produce the right output once
  • Whether you can explain your choices and their limits when asked about the submission later

How to prepare

  • Write one data transformation in Python or TypeScript with typed inputs, validation, and tests you write before the implementation
  • Practise a short README format: assumptions, what you chose not to handle, complexity, and what you would change with more time
  • Confirm with the recruiter whether AI assistants are permitted, and if they are, practise using one for scaffolding while writing the logic and tests yourself
PracHub interview research ↗
03

Core Interview Loop

reported

The source notes describe the core loop as several technical rounds plus a conversation with the hiring manager. Prepare for the hiring manager conversation separately from the technical rounds. The reported behavioral questions cover a critical technical tradeoff made under time pressure, a disagreement with a product direction or a teammate's architecture choice, prioritizing in an ambiguous environment, mentoring a junior engineer, and why you want to work at Cohere. Each needs a specific story with your own decisions in it. A story about what the team did gives the interviewer very little to go on.

What to demonstrate

  • Whether your stories show decisions you personally made, what you traded away, and what happened afterwards
  • Whether you settled a disagreement with evidence such as a benchmark, a prototype or a written proposal, rather than by seniority
  • Whether your reasons for wanting this role are specific and consistent with the work you describe

How to prepare

  • Write one story for each reported behavioral prompt as a timeline of decisions, then rewrite every sentence that starts with a team subject so it names your part
  • Prepare questions for the hiring manager about what the team is on call for, what shipped recently, and how work is split between new systems and existing ones
  • Run the stories past a peer who asks why at each step, and cut any claim you cannot back with a number or a date
PracHub interview research ↗
04

Live Coding

reported

Most of the reported coding questions are data-structure or algorithm problems with a language-model or pipeline setting: tokenizing streaming text under tight memory limits, a custom data structure for managing context windows, a graph traversal for dependency resolution, and optimizing an existing algorithm for high-throughput vector similarity search. The remaining one asks for clean, production-ready Python or TypeScript for a new data transformation, so practise writing readable, tested code as well as correct code. Write a correct simple version first, state its cost, then improve it while the working version stays on screen. Most of the mistakes in these problems are at the edges, such as a token split across two chunks, a dependency cycle, or an empty input, so trace those yourself before you say you are finished.

What to demonstrate

  • Whether the solution handles boundary cases without prompting: empty input, one element, duplicates, cycles, and data split across chunk boundaries
  • Whether the complexity you state matches the code, including memory when the problem sets a memory limit
  • Whether you explain the data structure choice in terms of the operations the problem needs

How to prepare

  • Solve streaming tokenization with a remainder buffer that carries partial tokens between chunks, and test a token split exactly at the boundary
  • Implement topological sort with Kahn's algorithm, report that a cycle exists when nodes are left unprocessed, and state O(V + E)
  • Implement a randomized set with O(1) insert, delete and random pick using an array plus a hash map, removing by swapping with the last element
PracHub interview research ↗
05

System Design

reported

The reported design questions deal with serving AI workloads and moving data: a distributed system for high-volume generative AI API requests, real-time monitoring and logging across regions, an ingestion pipeline that connects outside data sources to a core platform, high availability and low latency during traffic spikes, and managing state, caching and rate limiting across microservices. Start by fixing the scope and the load, then name the unit of cost. For language-model traffic, a request count says little about the work involved. Spend most of your time on failure modes and trade-offs rather than on the list of components.

What to demonstrate

  • Whether you state the load, the cost unit and the latency target before choosing components
  • Whether you explain how the system behaves when a dependency fails, when traffic spikes, and when a region goes down
  • Whether each choice about caching, rate limiting or consistency comes with the trade-off it costs

How to prepare

  • Design a rate limiter for LLM chat traffic that limits by tokens and concurrent streams, and say what happens when the counter store is unreachable
  • Design an ingestion pipeline with idempotent writes, capped exponential backoff with jitter, a dead-letter queue and per-source cursors
  • Compare event-driven and REST integration for a real-time AI feature, naming what each gives up on latency, ordering and back-pressure
PracHub interview research ↗
06

Distributed Systems Debugging

reported

The source notes describe an exercise in which interviewers add new findings while you work through a simulated production issue. Practise it as a method rather than trying to recall answers. Keep a visible list of hypotheses, ask for the specific metric, log or trace that would confirm or rule out each one, and when a new finding arrives, say which hypotheses it eliminates before you suggest a fix. The reported debugging questions include cascading failure under heavy load, intermittent memory leaks, database deadlocks diagnosed from logs and metrics, network partitions in a globally distributed cluster, and a slow query in a high-throughput data store.

What to demonstrate

  • Whether you form testable hypotheses and ask for evidence that separates them, rather than guessing a fix
  • Whether you revise your view openly when the interviewer introduces a new finding
  • Whether you separate short-term mitigation from root cause, and say which one you are working on
  • Whether your reasoning is clear enough out loud for the interviewer to follow and redirect

How to prepare

  • Have a partner run an incident from written notes and reveal one finding at a time while you narrate your hypotheses and requests
  • Write a triage checklist covering what changed recently, the blast radius, latency, errors, traffic and saturation per service, and which dependency failed first
  • For each reported scenario (memory leak, cascading failure, deadlock, partition, slow query), write the first three signals you would check and what each would rule out
PracHub interview research ↗

2 candidate reports. Individual accounts describe a particular role and hiring cycle.

Software Engineer

Cohere Software Engineer Interview Experience — Rejected at the ML System Design Round

Online Assessment → HR Screen → Technical ScreenOutcome: rejected

I interviewed for a role on the agent platform team. The process was: OA -> HM call (behavioral) -> system design -> VO (two rounds of debugging and coding). The OA was basic Python string parsing. The only tricky part was the last question, where the argument you're passed is callable — you need to print the function's name, its input/output format, and the expected result. You need to know whic…

Read full experience

PracHub editorial advice for the preparation topics above.

01

Proposing a fix in the debugging simulation before the evidence supports it

When the interviewer reveals findings as you go, a fix based on the first symptom (restart the pods, add memory, scale out) is usually overturned by the next finding. Keep a short, visible list of hypotheses and say which metric, log line or trace would confirm or rule out each one. Ask for that specific evidence. When a new finding arrives, say which hypotheses it eliminates before you propose anything. Mitigate first (shed load, roll back, fail over), then look for the root cause, and say which of the two you are doing.

02

Losing tokens that span a chunk boundary in the streaming-text coding question

The reported tokenization question sets a tight memory limit, so reading the whole stream into one string solves a different problem. Keep a small remainder buffer that holds the incomplete token at the end of each chunk, add it to the front of the next chunk, and flush it at the end of the stream. Before you call it done, test a token split exactly at the boundary, a chunk with no delimiter, an empty chunk, and a multi-byte character split across chunks if the input arrives as bytes. State the memory bound: the longest token plus one chunk.

03

Rate limiting generative AI traffic by request count alone

In the reported designs for generative AI APIs and rate limiting, a requests-per-second limit treats a short prompt and a long generation as the same cost. Say which unit you limit (requests, tokens, concurrent streams), where the counter lives, and whether admission fails open or closed when the counter store is unreachable, and why. Cover streaming responses that keep connections open, return 429 with Retry-After, and explain how limits stay roughly consistent across gateway instances and regions.

04

Retrying calls to external sources without idempotency or backoff

Ingestion and integration questions come down to what happens when an external source rate-limits you or times out. A tight retry loop makes the outage worse, and a blind retry creates duplicate records. Give each record a stable key so every write is an idempotent upsert, retry with capped exponential backoff and jitter, honor Retry-After, send poison records to a dead-letter queue with the error attached, and keep a per-source cursor so a restart resumes where it stopped instead of re-reading everything.

05

Submitting take-home or tool-assisted code you cannot defend line by line

The source notes say tools such as Claude Code may be allowed in specific technical evaluations. Confirm the rules for each stage with your recruiter instead of assuming. Where a tool is allowed, use it for scaffolding and boilerplate, and write the logic, edge cases and tests yourself. Be ready to explain every line, why you chose each structure, and what your tests do and do not cover. A short README of assumptions and trade-offs is easier to defend in a follow-up than clever code you cannot walk through.

Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.

12 technical prompts3 include a worked solution

Solve a complex graph traversal problem related to dependency resoluti…

medium
data structures and algorithms

Solve a complex graph traversal problem related to dependency resolution in distributed pipelines.

Approach
  1. Choose the data structure from the access pattern, not from familiarity.
  2. State the target complexity and say which constraint rules the naive version out.
  3. Restate the input: its shape, its size, and what is guaranteed about it.
Follow-up
  • Which test case would catch an off-by-one here?
  • What is the worst case, and how likely is it on real data?

Write clean, production-ready code in Python or TypeScript to solve a …

medium
data structures and algorithms

Write clean, production-ready code in Python or TypeScript to solve a novel data transformation challenge.

Approach
  1. State the target complexity and say which constraint rules the naive version out.
  2. Walk one small example through your approach before writing the whole thing.
  3. Choose the data structure from the access pattern, not from familiarity.
Follow-up
  • What is the worst case, and how likely is it on real data?
  • How does this change if the input no longer fits in memory?

Implement a custom data structure to efficiently manage and retrieve c…

medium
data structures and algorithms

Implement a custom data structure to efficiently manage and retrieve context windows for language models.

Approach
  1. Choose the data structure from the access pattern, not from familiarity.
  2. State the target complexity and say which constraint rules the naive version out.
  3. Name the brute-force solution and its complexity before improving on it.
Follow-up
  • What is the worst case, and how likely is it on real data?
  • Which test case would catch an off-by-one here?

Write a function to process and tokenize streaming text data under tig…

medium
data structures and algorithms

Write a function to process and tokenize streaming text data under tight memory constraints.

Approach
  1. Name the brute-force solution and its complexity before improving on it.
  2. State the target complexity and say which constraint rules the naive version out.
  3. Choose the data structure from the access pattern, not from familiarity.
Follow-up
  • How does this change if the input no longer fits in memory?
  • Which test case would catch an off-by-one here?

Fold a deduplicated usage stream into hourly rollups

easyWorked solution
aggregationdeduplicationwatermarksexact-arithmetic

You are given one day of usage_event rows, up to 250 million, each carrying event_id, tenant_id, workspace_id, environment, sku, quantity numeric(20,6), idempotency_key, occurred_at and ingested_at. Produce usage_rollup_hourly cells keyed (tenant_id, workspace_id, sku, hour_start) with quantity_sum, event_count and source_max_ingested_at. An event counts once per (tenant_id, idempotency_key). The rollup grain has no environment column, so state your filter. One pass. Give your time and space bounds, and say what the deduplication actually costs in memory.

Approach
  1. Bucket on occurred_at, never ingested_at: hour_start = date_trunc('hour', occurred_at at time zone 'UTC'). The two columns answer different questions. occurred_at says which hour the customer is billed for; ingested_at says how current the fold is. Using the second for the first makes late data invisible instead of correctable.
  2. The fold is trivial and the deduplication is the entire cost, so price it before designing anything clever. An exact set over (tenant_id, idempotency_key) at 250M entries, stored as a 16-byte 128-bit hash in an open-addressed table at 0.7 load factor, needs about 357M slots at 16 bytes each, roughly 5.7 GB. The fix is partitioning by hash(tenant_id) % P so each shard holds 1/P of the set and no tenant's keys straddle shards.
  3. Rule out a Bloom filter as a replacement, in the right direction: a false positive reports 'already seen' for an event never seen, so you drop a real event and lose revenue with no error raised. It is usable only as a negative pre-filter in front of the exact set, where a miss is conclusive and a hit must fall through to the real lookup.
  4. Accumulate in scaled integers, not binary floating point. numeric(20,6) admits values below 10^14, so one event scaled to micro-units can reach 10^20, past int64's 9.22 x 10^18; use a 128-bit or arbitrary-precision accumulator unless you first bound the per-event maximum. binary64 represents integers exactly only to 2^53, about 9.01 x 10^15, and cannot represent 0.1 at all, so two runs that sum in different orders disagree.
  5. Carry source_max_ingested_at = max(ingested_at) over the events folded into each cell, and count event_count over accepted, post-dedup events. Without that watermark there is no way to prove later what a number did and did not include, which is the first question any reconciliation asks.
  6. State the environment filter explicitly, because the rollup grain cannot record it. A fold that quietly includes staging bills non-production traffic; one that quietly excludes it loses a cost signal. Production-only is the billing answer, and either way it belongs in the job name and the output metadata. Complexity: O(n) time, O(distinct dedup keys) space, dominated by the dedup set rather than by the cells.
Worked solution 25 min
  1. Write both key tuples down before any code: dedup key (tenant_id, idempotency_key), cell key (tenant_id, workspace_id, sku, hour_start), with hour_start derived from occurred_at in UTC.
  2. Build a 10,000-row fixture containing one event duplicated three times under the same idempotency_key, two events sharing an idempotency_key across different tenant_id values, one event whose occurred_at is two hours before its ingested_at, and one staging event inside an otherwise production cell.
  3. Fold it and assert each of those four expectations separately rather than eyeballing a grand total.
  4. Re-run with the input shuffled and diff the output files.
  5. Size the dedup set for 250M keys using the load-factor arithmetic and write the number down next to the fixture.
EXPECTED RESULTThe triplicate contributes one event and its quantity once. The two same-key, different-tenant events both count, because the dedup key is the pair. The late event lands in the hour of its `occurred_at` while that cell's `source_max_ingested_at` advances to the later timestamp. The `staging` event is included or excluded per the stated filter and never silently.
Follow-up
  • A producer retries at 23:59:59 and the retry lands at 00:00:01. The unique index on the daily-partitioned table must include the partition key. What gets double-counted, and what is the smallest change that fixes it?
  • The consumer acknowledges its batch before committing the fold. Which failure loses revenue now, and which arrangement duplicates instead?
  • What makes a re-run over the same day produce byte-identical rollups?

Day one measures instead of guessing, under a fixed rubric, and the remaining hours are allocated in proportion to the gaps before any studying begins. The allocation is deliberately not renegotiated midweek, because the area that feels worst on day three is usually the one that is moving.

Small steps. Visible outcomes.0 / 7 completed
ONE WEEK · YOUR PACE

Prepare, practise & reflect

One practical outcome each day. Spend longer where you need it.

0 / 7 done
01Map the six rounds and prepare the recruiter screen
  • List the six stages candidates report (Recruiter Screening, Technical Assessment, Core Interview Loop, Live Coding, System Design, Distributed Systems Debugging). Next to each, write the question categories this guide links to it.
  • Attempt the reported dependency-resolution graph question cold, and score yourself on correctness, edge cases and how clearly you explained your approach.
  • Write your recruiter questions: take-home or live assessment, accepted languages, whether AI coding tools are allowed at any stage, and what the debugging exercise looks like.
  • Prepare a short spoken background summary that ties your work to distributed systems, platform work, developer products or data integrations.

Deliverable: A one-page loop map with your weakest round circled, plus a written list of recruiter questions.

Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗
02Live coding: streams and data structures
  • Solve streaming tokenization under a memory limit using a remainder buffer, then test a token split at the boundary, an empty chunk and a chunk with no delimiter.
  • Design a context-window structure that stays under a token budget by evicting the oldest entries, and state the cost of each operation.
  • Implement a randomized set with O(1) insert, delete and uniform random pick using an array and a hash map.
  • Before running any solution, write down what it returns for empty input, one element and all duplicates, then check your predictions.

Deliverable: Three working solutions, each with the edge cases you traced written above it.

Practice prompt ↗Practice prompt ↗
03Coding: graphs, intervals and production-ready transformations
  • Implement dependency resolution with Kahn's algorithm, report that a cycle exists when nodes are left unprocessed, and state O(V + E). As an extension, run a DFS over the leftover subgraph to extract one actual cycle, since the leftover set also contains nodes downstream of the cycle.
  • Solve maximum overlapping events with a sweep line, and state whether an event ending and another starting at the same timestamp count as overlapping.
  • Work through the guide's worked coding exercise on folding a deduplicated stream into hourly rollups, including the shuffled-input check.
  • Write one data transformation in Python or TypeScript with validation and tests, to the standard of the reported production-ready code question.

Deliverable: A tested transformation module and a written complexity note for each problem.

Practice prompt ↗Practice prompt ↗
04System design: serving generative AI APIs
  • Design a system for high-volume generative AI API requests: gateway, authentication, admission control, queueing, streaming responses and autoscaling.
  • Add a rate limiter that limits by tokens and concurrent streams, and decide whether it fails open or closed when the counter store is unreachable.
  • Design a cache for LLM reads: what can be cached, the cache key, how stale an entry may be, and what happens on a cache stampede.
  • Work through the guide's worked design exercise on a machine-readable error contract, including the status-to-retry table.

Deliverable: One written design covering the interface, data flow, a capacity estimate and a table of failure modes.

Practice prompt ↗Practice prompt ↗Worked solution ↗
05System design and SQL: ingestion, observability and data stores
  • Design an ingestion pipeline for external sources with idempotent upserts, capped backoff with jitter, a dead-letter queue and per-source cursors.
  • Compare event-driven and REST integration for a real-time AI feature on latency, ordering and back-pressure.
  • Sketch real-time monitoring for multi-region AI services: metrics, logs and traces, where they are aggregated, and which alerts page someone.
  • Work through the guide's worked SQL exercise on keyset pagination, then read the EXPLAIN plan of a slow query you have written and name the missing index.

Deliverable: A pipeline diagram with a failure table, and one query plan annotated before and after the fix.

Practice prompt ↗Practice prompt ↗
06Distributed systems debugging drill
  • Have a partner run an incident from written notes and reveal one finding at a time while you narrate your hypotheses and ask for specific evidence.
  • Run at least two of the reported scenarios: cascading failure under load, an intermittent memory leak, database deadlocks from logs and metrics, a network partition, or a slow query in a high-throughput store.
  • Write a triage checklist: recent changes, blast radius, latency, errors, traffic and saturation for each service, and which dependency failed first.
  • For each drill, write down separately what you did to mitigate and what the root cause was.

Deliverable: Two recorded incident drills and a one-page triage checklist.

Practice prompt ↗Practice prompt ↗
07Core loop stories and a full mock
  • Prepare one story for each reported behavioral prompt: a tradeoff made under time pressure, a disagreement, prioritizing in an ambiguous situation, mentoring a junior engineer, and why Cohere.
  • Rewrite each story so every decision names who made it, and check that team size, timeline and your role stay the same across all of them.
  • Prepare questions for the hiring manager about on-call, recent projects and how work is split between new and existing systems.
  • Run a mock with one coding problem, one design prompt and one debugging scenario, with the interviewer briefed to interrupt and add new information.

Deliverable: A story sheet with one entry per behavioral prompt, and written notes from the mock naming the round you would most likely lose.

Practice prompt ↗Practice prompt ↗Worked solution ↗

Expand any day for tasks and deliverables. Your progress is saved on this device.

The first five prompts below are the reported behavioral questions, covering tradeoffs, disagreement, prioritization, mentoring and motivation; the last is one of this guide's own practice drills. For each one, prepare a specific story with your own decisions in it: what you knew at the time, the option you rejected, what you measured, and what you would change now. If a story only says what the team achieved, rewrite it to show your part.

Describe a situation where you disagreed with a product direction or a…

medium
behavioural and engineering judgement

Describe a situation where you disagreed with a product direction or a teammate's architectural choice; how did you resolve it?

Approach
  1. State the situation in two sentences and spend the rest on the reasoning.
  2. Close with what you would do differently, concretely.
  3. Give the blast radius: what could have broken, and what you measured.
Follow-up
  • How did you know your change caused the improvement?
  • What did you decide not to do, and why?

Resolve a review disagreement over a quota check

easy
code reviewisolation levelswrite skewdisagreement

A colleague's pull request enforces a per-tenant quota by selecting the current count and then inserting when it is under the limit. You flag it as a race. They reply that the transaction already runs at repeatable read, so the snapshot makes it safe, and the tests pass. Walk through taking that disagreement to a resolution: what you write in the review, what you demonstrate rather than assert, which fix you propose and why, and what you do if they still disagree after all of it.

Approach
  1. Answer the claim precisely instead of restating your objection, because they have made a specific technical argument. In PostgreSQL, repeatable read is snapshot isolation; this is write skew, which snapshot isolation permits by design. Both transactions read a count that is stable within their own snapshot, insert disjoint rows that the other cannot see, and both commit, so the limit is exceeded by exactly the concurrency.
  2. Demonstrate rather than cite. Two psql sessions, both BEGIN ISOLATION LEVEL REPEATABLE READ, both select the count, both insert, both commit: it succeeds. Repeat at SERIALIZABLE and the second commit fails with serialization_failure, SQLSTATE 40001. That takes two minutes, ends the argument without anyone conceding a position, and leaves an artefact for the next reviewer.
  3. Offer the options with their costs rather than a verdict. Serialisable plus a retry loop on 40001 is correct but obliges every caller to retry and degrades under contention. An increment-and-compare on a counter row — update tenant_quota set used = used + 1 where tenant_id = $1 and used < limit returning used — is safe even at read committed, because a blocked updater re-evaluates the WHERE clause against the row version it finally locks, and zero rows returned means full. A unique or exclusion constraint that makes the surplus write fail is the third.
  4. Name the plausible non-fix explicitly, since it is what usually gets merged instead: folding the count into the insert as insert ... select ... where (select count(*) ...) < limit is still racy under read committed, because the subquery cannot see the other transaction's uncommitted rows. It looks atomic and is not.
  5. Say what you do if they still disagree: escalate the decision rather than the disagreement. Attach the reproduction, hand it to the service owner or a third reviewer, and state that you will not block the merge if the owner accepts the risk knowingly — and that you want that acceptance written down.
  6. Close with the general lesson worth leaving in the review thread: a passing suite is weak evidence for a concurrency claim because it runs one request at a time. Ask for a test that runs two.
Follow-up
  • Write the counter-row version. Does your answer change if the quota counts child rows rather than a column?
  • Under serialisable, who performs the retry, and what does the API client see if the retry also fails?
  • This is the third disagreement with the same reviewer this month. What changes in how you review?

Estimate a tenant-leading index migration you have never run

hard
estimationonline migrationindex buildsuncertainty

Someone needs a date. usage_event carries an index on (occurred_at) and needs (tenant_id, occurred_at); the largest tenant holds roughly a hundred times the median tenant's rows, the table is partitioned daily with years of retention, and you have never run a migration on a table this large. Give an estimate you would defend: how you decompose the work, the two or three numbers you would go and measure first, the range and confidence you state, and what you commit to when the person asking needs a single date today.

Approach
  1. Refuse the bare number and then give one anyway, in the form that is actually useful: a range plus the measurement that collapses it. 'Four to eleven days; one afternoon building this index on a restored copy of the largest partition takes that to within a day' is an answer, while 'it depends' is not.
  2. Decompose by failure mode rather than into equal chunks, because that is where estimates go wrong. On a partitioned parent you create the index ON ONLY the parent, build each partition's index with CREATE INDEX CONCURRENTLY, then ALTER INDEX ... ATTACH PARTITION, at which point the parent index becomes valid. CONCURRENTLY does not block writes but scans each partition twice, waits out older transactions, cannot run inside a transaction block, and on failure leaves an invalid index you must drop concurrently and retry.
  3. Name the two unknowns that dominate and price them: build time on one restored partition of realistic size, and whether the planner actually chooses the new index for the skewed tenant, since selectivity for a tenant holding most of the rows is a different question from selectivity for the median tenant. Both are half-day measurements against a replica, and both are cheaper than being wrong by a week.
  4. State the assumptions the range is conditional on, because that is what makes a slip a re-estimate instead of a credibility event: no partition above a stated row count, one concurrent build at a time so it does not compete with ingest for I/O, and an ingest backlog that can absorb the added write amplification while both indexes exist.
  5. Budget the step nobody budgets: verification and the old index's removal. Dropping the old index is fast, but deciding it is safe to drop means confirming no plan still uses it, and that confirmation waits on real traffic across a full weekly cycle rather than on your patience.
  6. Answer the single-date request honestly. Commit to a date for the first checkpoint — the measured build number from the replica — and to re-estimating on that date, and say plainly what you are not committing to yet. A date with a scheduled re-estimate is worth more to the asker than a confident wrong one, and you should say why in those words.
Follow-up
  • The concurrent build fails half way through the largest partition. What is the state of the database and what do you do next?
  • Your estimate slips by sixty percent. Which assumption broke, and at what point would you have known?
  • The person asking needs the date for a customer commitment. Does your answer change?
  • 01

    Tell me about a time you had to make a critical technical tradeoff under severe time pressure.

  • 02

    Describe a situation where you disagreed with a product direction or a teammate's architectural choice; how did you resolve it?

  • 03

    How do you prioritize tasks and maintain high quality standards in a fast-paced, ambiguous startup environment?

  • 04

    Share an example of a time you mentored a junior engineer or elevated the technical bar on your team.

  • 05

    Why do you want to build your career at Cohere, and how do your values align with its mission?

  • 06

    A colleague's pull request enforces a per-tenant quota by selecting the current count and then inserting when it is under the limit. They say repeatable read makes it safe and the tests pass. How do you take that disagreement to a resolution?

PracHub interview preparation framework ↗
Is this an official Cohere interview guide?

No. It is PracHub's own research and practice material for the Software Engineer role at Cohere. Rounds and questions reflect what candidates have reported, not a process Cohere has published, and they change over time. Confirm the current format and scope with your recruiter.

PracHub interview research ↗
What rounds do candidates report for the Cohere Software Engineer role?

Candidates report six stages: a recruiter screen, a practical technical assessment or take-home assignment, a core interview loop with several technical rounds and a hiring manager conversation, live coding, system design, and a distributed systems debugging exercise. Confirm with your recruiter which stages apply to your opening.

PracHub Software Engineer practice ↗
What happens in the distributed systems debugging round?

The source notes describe a simulated production issue in which interviewers add new findings while you work through it. Practise it as a method: keep a list of hypotheses, ask for the evidence that would separate them, say which hypotheses each new finding rules out, and keep mitigation separate from root cause. The reported debugging questions cover cascading failures, memory leaks, database deadlocks, network partitions and slow queries.

PracHub Software Engineer practice ↗
Which programming language should I use?

One reported coding question asks for production-ready code in Python or TypeScript. Beyond that, use the language you are most fluent in for live coding unless your recruiter says otherwise, and check which languages the technical assessment accepts before you start it.

PracHub Software Engineer practice ↗
Can I use AI coding assistants during the interviews?

The source notes say tools such as Claude Code may be permitted in specific technical evaluations. Do not assume this applies to every stage, and ask your recruiter which stages allow it. Where it is allowed, use the tool for scaffolding and write the logic and tests yourself, so you can explain every line if asked.

PracHub Software Engineer practice ↗
What system design topics should I prepare?

The reported design questions cover serving high-volume generative AI API requests, real-time monitoring and logging across regions, an ingestion pipeline that connects external data sources to a core platform, high availability during traffic spikes, and state, caching and rate limiting across microservices. An ML system design question on an enterprise research assistant with verifiable citations also appears in the question bank.

PracHub Software Engineer practice ↗
How should I prepare for the hiring manager conversation?

The source notes place a hiring manager conversation inside the core loop. Prepare stories where your own decisions are clear, covering tradeoffs, disagreements and mentoring, plus a specific answer to why this role. Bring questions about what the team is on call for, what it shipped recently, and how much of the work is new systems versus existing ones.

PracHub Software Engineer practice ↗
Sources & methodology 3 sources ↗

Official role evidence, timestamped platform data and clearly labeled preparation advice.