Together AI's Software Engineer role works on the AI Acceleration Cloud, infrastructure that virtualizes machine learning hardware, including NVIDIA GB200/GB300 GPUs and BlueField DPUs, for LLM inference and training. Customers use it to provision on-demand compute, managed Kubernetes clusters and Slurm workloads, and the same platform serves Together AI's internal products.
The role is described as sitting on the Together Cloud Platform or Together Cloud Infrastructure team. The listed work covers backend services and API microservices behind customer-facing products, written primarily in Go; infrastructure as code for physical and virtual resources; design and code reviews, developer documentation and testing strategy; and an on-call rotation for infrastructure incidents. The systems involve Infiniband networking, exabyte-scale data pipelines and distributed scheduling.
The listed must-haves are 5+ years building fault-tolerant distributed systems and API microservices, strong proficiency in a backend language with Go preferred, systems knowledge across compute, networking and storage (concurrency, memory management, performant I/O), hands-on PostgreSQL, and Kubernetes, containers and CI/CD. Nice-to-haves include QEMU/KVM and KubeVirt, SR-IOV, GPU virtualization, Infiniband or RDMA, DPUs and SmartNICs, CUDA or NCCL, and Kafka, Airflow or Kinesis.
The reported questions track that list closely: distributed scheduling and storage design, a Kubernetes operator in Go, GPU virtualization trade-offs, data-center networking, Terraform and Ansible delivery, Prometheus observability, and PostgreSQL and Linux performance debugging. Be ready to explain how each tool works underneath, not only when you would choose it. The company's open-source work includes FlashAttention and RedPajama, so a quick read on what each one is makes useful background.
Recruiter Screen
reportedCandidates describe the first stage as a recruiter conversation to align on background, career goals and compensation expectations. Use it to shape the rest of the process. Connect your experience to the areas the role lists: distributed systems and API microservices, Go, compute, networking and storage fundamentals, PostgreSQL and Kubernetes. Then find out which of the reported technical focuses your phone screen will use: systems programming, concurrent coding or architecture design. The role is described as sitting on either the Together Cloud Platform or the Together Cloud Infrastructure team, so ask which one you are being considered for. The answer tells you whether to weight service design or hardware-level infrastructure.
What to demonstrate
- Whether your background lines up with the listed must-haves: fault-tolerant distributed systems, a backend language (Go preferred), systems knowledge across compute, networking and storage, PostgreSQL and Kubernetes
- Whether your career goals fit infrastructure work that sits between physical hardware and cloud software
- Whether your compensation expectations fit the role, a topic candidates report comes up at this stage
How to prepare
- Write two short project summaries that map to the listed requirements: one distributed system you built or ran in production, and one piece of infrastructure or low-level work such as virtualization, networking, storage or a Kubernetes controller
- Settle a compensation range before the call so you can answer with a number
- Ask which team the process is for, which languages the coding screens accept, and whether the phone screen is coding, systems programming or architecture design
Technical Phone Screens
reportedCandidates report one or more technical phone screens focused on systems programming, concurrent coding or high-level architecture design. The format can be any of the three, so prepare all three to screen depth rather than betting on one. For code, reported guidance stresses clean, idiomatic Go with attention to concurrency and tests. A goroutine-based solution that shuts down cleanly and passes the race detector covers much of that. For systems programming, explain what the runtime and kernel actually do, not only which API you would call. For architecture, a compact design with a clear failure story is stronger than a broad one with no failure handling.
What to demonstrate
- Whether you can write correct concurrent code: bounded goroutines, channels or mutexes chosen for a stated reason, cancellation through context, and no leaked goroutines
- Whether you can explain the systems behaviour underneath the code, such as garbage-collection cost, buffer reuse and the file I/O path
- Whether a design answer states its failure modes and delivery semantics, not only its components
How to prepare
- Implement a bounded worker pool in Go with context cancellation and error propagation, then run it under go test -race
- Practise the coding titles in the question bank: detecting and breaking cycles in pod dependencies, scheduling GPU pods and draining a node, splitting a chunked text stream into line-balanced parts, and the first unique character index
- Prepare spoken answers on GC versus non-GC memory management and on file I/O bottlenecks, including zero-copy and asynchronous system calls
- Rehearse one scheduling design in compact form that ends with what happens during a network partition
Virtual or On-site Interview Loop
reportedCandidates report a virtual or on-site loop covering distributed systems, networking, infrastructure automation and behavioural alignment. Reported guidance says design discussion goes down to the hardware and network layers: how memory is managed, how packets are routed, how database transactions are isolated. The loop spans several areas, so prepare each one separately: a distributed scheduling or storage design, a networking and virtualization comparison, an infrastructure-as-code or observability pipeline, and behavioural stories about outages and disagreements. Keep the facts about any project you mention identical in every conversation.
What to demonstrate
- Whether a distributed design handles node failure, network splits and hardware degradation, with its delivery and consistency guarantees stated
- Whether you understand networking and virtualization choices such as VLAN, VXLAN and VPC isolation, Infiniband and RDMA, and hardware passthrough
- Whether infrastructure automation is safe to run against production: reviewed plans, staged rollout, drift handling and guarded node lifecycle actions
- Whether behavioural stories show ownership of incidents and a clear, evidence-based way of resolving technical disagreement
How to prepare
- Work the reported design questions in this guide (at-least-once task scheduling under partitions, a global management plane, on-demand GPU scheduling, multi-exabyte object storage), taking each one layer below the box diagram
- Write one-paragraph comparisons of VLAN vs VXLAN vs VPC and of VFIO vs SR-IOV vs PCIe passthrough, and practise defending them aloud
- Sketch a Terraform and Ansible delivery pipeline and a Prometheus/Grafana node-health pipeline that cordons and drains nodes, naming the safety check in each
- Prepare outage, disagreement and shifting-requirements design-doc stories with a fact sheet per project so figures stay the same in every round
PracHub editorial advice for the preparation topics above.
Claiming exactly-once execution in the at-least-once task scheduler design
At-least-once means a task can run twice. A worker can finish and lose its acknowledgement, or it can be partitioned away while its lease expires and another worker picks the task up. Say that directly, then make it safe: leases with expiry, a fencing token checked on every commit so a returning stale worker is rejected, and task execution that is idempotent on the task id. Walk through a partition and its healing before the interviewer asks for it.
Treating on-demand GPU scheduling as generic bin-packing
Counting free GPUs per node misses what makes the problem hard. Multi-GPU jobs need all their devices at once (gang allocation), placement should respect interconnect and network topology, and scattering small jobs strands capacity that no large job can use. Define a fragmentation metric, say when you preempt or migrate work, and explain how you drain a node without killing long-running training jobs.
Naming Kubernetes, etcd or a hypervisor without explaining how it works
Reported guidance asks candidates to explain tools under the hood. For the custom operator question, describe the reconcile loop as level-triggered and cover the informer cache and the workqueue. Explain why the workqueue never hands the same key to two workers at once, what raising MaxConcurrentReconciles changes, and why reconcile must be idempotent, since it will run again against a stale cache. Bring the same depth to VFIO, SR-IOV and PCIe passthrough.
Writing Go concurrency that passes once but leaks goroutines or races
Unbounded goroutine spawning, channels nobody closes, sends that block forever after the receiver has returned, and shared maps without a lock can all pass a single happy-path run. Bound concurrency with a worker count or semaphore, pass a context for cancellation, make it explicit which side closes each channel, and say you would run the tests with -race. Explain these choices aloud, because Go-specific concurrency and performance questions are reported in systems-focused rounds.
Guessing at a cause in a performance-debugging question
For a throughput drop in PostgreSQL, CPU, disk or network, start with a method, not a hypothesis: check utilisation, saturation and errors for each resource, then narrow down. Name the tool at each step: top or pidstat and perf for CPU, iostat -x for disk queueing and latency, ss and interface counters for the network, and pg_stat_activity wait events with pg_blocking_pids() for database lock waits. For each step, also say what reading would rule that cause out.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
How does memory management differ when writing high-throughput network…
How does memory management differ when writing high-throughput network applications in a garbage-collected language like Golang versus a non-garbage-collected language?
Approach
- Distinguish a value from a reference to it, and say which one you handed out.
- Reach for the cheapest primitive that closes the race, not the broadest lock.
- Say what the runtime actually does before reasoning about the code.
Follow-up
- How would you prove the race exists rather than suspect it?
- Where could this allocate more than you expect?
Describe the performance differences between VFIO, SR-IOV, and standar…
Describe the performance differences between VFIO, SR-IOV, and standard PCIe passthrough when virtualizing hardware accelerators like GPUs.
Approach
- Say what the runtime actually does before reasoning about the code.
- Reach for the cheapest primitive that closes the race, not the broadest lock.
- Name what is shared across threads and what owns each piece of state.
Follow-up
- How would you prove the race exists rather than suspect it?
- Where could this allocate more than you expect?
Seal an hour under late data with bounded memory
Metering ingest reads 256 partitions at 10,000 to 40,000 events/second. Events carry occurred_at and ingested_at, and during a producer replay the gap between them is hours. Seal each UTC hour once no more than 50 parts per million of that hour's eventual quantity can still arrive, using memory that does not grow with the size of the replay. Define the watermark, the lateness parameter and how you measure it, the structure holding open hours, and the write that performs the seal. State what an idle partition does to your watermark.
Approach
- Two clocks, two jobs. Bucket by
occurred_at, because that is the hour the customer is billed for, and advance the watermark oningested_at, because that is what the fold has consumed and whatsource_max_ingested_atrecords. Conflating them is what makes late data invisible. - The global watermark is the min over partitions of each partition's committed
ingested_at, not the max: the fold is trustworthy only as far as the slowest partition. The consequence is that one idle partition pins the watermark forever and nothing seals, so an idle partition must promote its watermark to wall clock after a stated idle timeout, and that timeout becomes a correctness parameter, because a partition that is slow rather than idle gets sealed past. - Choose the lateness L from the measured distribution of
ingested_at - occurred_at, weighted by quantity rather than by event count. The target is 50 ppm of the hour's quantity, and a replay is rare in events while carrying disproportionate mass, so an event-weighted quantile picks an L that is comfortably wrong at exactly the moment it matters. - Measure that quantile in bounded memory. A Greenwald-Khanna summary gives epsilon-approximate quantiles in O((1/epsilon) log(epsilon n)) space; a t-digest costs more per merge but has relative error that tightens at the tails, which is the half of the distribution you are reading at p99.99. Keep a separate summary per tenant class, because one tenant's batch importer is not the population.
- Hold open hours in a min-heap keyed by
hour_start. When the watermark advances, pop every hour withhour_end + L < Wand seal it: O(log H_open) per advance and O(1) amortised per event to touch its bucket. Memory is open hours multiplied by distinct(tenant, workspace, sku)keys, so cap the number of simultaneously open hours and spill the oldest intousage_rollup_hourlyasstatus='open'with arevisionbump. While an hour is open the row is upsertable, so the store is your overflow. - The seal itself is a conditional write:
update ... set status='sealed', sealed_at=now() where status='open' returning .... Two sealers race on every restart, and the loser must see zero rows and stop rather than write a second value. After the seal, an event for that hour is not an upsert but an adjustment, andsource_max_ingested_atis what proves it arrived afterwards.
Worked solution 40 min
- Replay a day of events with a synthetic lateness distribution: 99.9% under two minutes, plus a 0.05% tail at four to six hours that carries 3% of total quantity.
- Compute the p99.99 lateness two ways, event-weighted and quantity-weighted, and put the two numbers side by side.
- Implement the min-heap of open hours with the watermark as the min over 256 partitions, then stall one partition for 20 minutes and observe what seals.
- Set the idle-partition timeout to 60 seconds, repeat the stall, and measure how much quantity arrives after the seal.
- Attempt the seal from two workers at once and confirm the conditional update lets exactly one through.
Follow-up
- A replay starts during the sealing window for a period you are about to close. What do you do, and what is the customer-visible consequence of each option?
- Your measured quantity-weighted p99.99 lateness is six hours and the invoice must be issued at 02:00 UTC on the first. How do you reconcile those two numbers?
- How would you detect that L has drifted before it costs you an hour's quantity?
Decide which facts an invoice line copies instead of joining
invoice_line_item already denormalises tenant_id, which is reachable through invoice_id, and stores amount_minor even though quantity times unit_price_micros would recompute it. A reviewer asks you to normalise both away, and separately asks whether the tenant's legal name and billing address should be copied onto the invoice header. Decide each case. For every field you keep denormalised, name the read pattern or the invariant that justifies it, the anomaly the copy can develop, and the mechanism that prevents that anomaly here.
Approach
- Split the question into two kinds of copy, because they fail differently. A copy of a currently mutable fact is a cache: it drifts and needs invalidation. A copy of a fact frozen at write time is not a cache at all, it is the record of what happened, and normalising it away destroys information the source no longer holds.
- Keep tenant_id on the line. It costs 8 bytes, it leads every index on the table so no read is ever accidentally cross-tenant, and it turns a wrong join into an empty result rather than another tenant's money. Prevent the drift structurally: a unique constraint on invoice (invoice_id, tenant_id) plus a composite foreign key from the line on (invoice_id, tenant_id) makes a mismatched pair impossible, so the database enforces agreement instead of a code review.
- Keep amount_minor. Rounding must happen exactly once, at a named site, with a stated mode (half-even here). If readers recompute from quantity and unit_price_micros, every reader owns a rounding decision, and half-up and half-even diverge systematically across thousands of lines rather than cancelling out. A check constraint can bound the stored value but deliberately cannot re-derive it.
- Copy the legal name and billing address onto the invoice header, written once and never updated. The statement must show what was true when it was sealed, and the tenant record will change afterwards. This is a snapshot for the same reason
source_rollup_watermarkis stored per line: without it, nobody can reconstruct what the customer was told. - Name the read pattern that pays for all of it. Rendering, dispute response and export are per-tenant, per-period reads over thousands of lines that would otherwise join back to slowly changing dimensions that no longer hold the historical value. The write side is a once-per-period batch, so the extra columns cost nothing that matters.
- Concede the case where the reviewer is right: a mutable operational attribute such as the tenant's current plan name has no business on a line. If a report wants it, join. If a statement needs the plan as of the period, that is another snapshot and it belongs on the header with the rest.
Worked solution 25 min
- Write the DDL: unique (invoice_id, tenant_id) on invoice, the composite FK from the line, and a comment on each denormalised column saying whether it is a snapshot or a cache.
- Attempt to insert a line whose tenant_id differs from its invoice's and confirm the foreign key rejects it.
- Rename a tenant, re-render a sealed invoice, and confirm the rendered name is the one stored on the header.
- Recompute amount_minor from quantity times unit_price_micros for a thousand synthetic lines rounding half-up, sum both ways, and record the divergence from the stored half-even values.
Follow-up
- Write the composite foreign key and the unique constraint it requires on the parent. What does it cost on every line insert, and what does it do to a bulk load?
- A tenant is renamed after being invoiced. Which rows change, and what does the customer see on last quarter's PDF?
- Where does currency live, and what breaks if a tenant's billing currency changes between two periods?
Find the join that inflates every invoice total
invoice_line_item holds line_id, invoice_id, tenant_id, sku, rate_tier, quantity, unit_price_micros, amount_minor (bigint), currency, kind, voided_at. invoice_payment_attempt holds attempt_id, invoice_id, tenant_id, amount_minor, status (succeeded, failed, pending), created_at, and an invoice has many attempts. A finance report runs select i.invoice_id, sum(l.amount_minor), count(p.attempt_id) from invoice i join invoice_line_item l using (invoice_id) join invoice_payment_attempt p using (invoice_id) group by 1 and the totals are wrong. Say precisely what the sum now equals, and write a version that is also correct for invoices with zero attempts.
Approach
- Compute what the query actually returns before fixing it. The two joins form a Cartesian product per invoice, so each line row repeats once per attempt row:
sum(l.amount_minor)is the true total multiplied by the attempt count, andcount(p.attempt_id)is attempts times lines. Three lines and two attempts report double the money and six attempts. - Reject the reflex repair.
count(distinct p.attempt_id)does fix the count, because attempt_id is unique.sum(distinct l.amount_minor)does not fix the sum, because two legitimate lines with equal amounts collapse into one. DISTINCT inside an aggregate deduplicates values, not rows, and the difference stays invisible until two lines happen to match. - Aggregate each branch to invoice grain before joining: one CTE summing lines by invoice_id, one counting attempts by invoice_id, then join the two results. A LATERAL subquery per invoice is equivalent and sometimes plans better when the outer set is small. Either way every aggregate stays at the grain it was defined at.
- Keep invoices with no attempts by making the attempt branch a LEFT JOIN with
coalesce(attempt_count, 0). An inner join here silently drops every unpaid invoice, which is usually the exact population finance is asking about. - Push each filter to its own grain:
where l.voided_at is nullbelongs inside the line CTE, not the outer query, or it would also filter the attempt branch through the join. Put the tenant predicate on both branches, since the denormalised tenant_id is what stops a wrong join crossing tenants. - Leave yourself a standing check: an invoice total is a function of its non-voided lines and of nothing about payments, so if changing the payment filter moves the money figure, the fan-out is back.
Follow-up
- Add a third branch for credit notes applied to the invoice. Does the CTE shape still hold, and when would a single pass with
filter (where ...)be better? - Over 500k invoices this report takes minutes. Which grain would you materialise, and how do you keep it correct when a line is voided?
- The same report is needed per tenant per month. What index makes the line CTE cheap?
How would you design a distributed, highly available task scheduling s…
How would you design a distributed, highly available task scheduling system that guarantees at-least-once execution under heavy network partitions?
Approach
- Name the failure you are designing for, then the recovery path.
- Name the read and write paths separately; they rarely have the same bottleneck.
- State the consistency you need, and where you are willing to be stale.
Follow-up
- How does this behave when that dependency is down for an hour?
- What breaks first when traffic grows ten times?
How would you set up and automate Infiniband partitioning and parallel…
How would you set up and automate Infiniband partitioning and parallel storage provisioning in a bare-metal data center?
Approach
- Name the read and write paths separately; they rarely have the same bottleneck.
- State the consistency you need, and where you are willing to be stale.
- Choose a partition key and say what query it makes expensive.
Follow-up
- What breaks first when traffic grows ten times?
- How does this behave when that dependency is down for an hour?
Explain how you would architect a global management plane to control c…
Explain how you would architect a global management plane to control compute, networking, and storage across dozens of independent data centers.
Approach
- Choose a partition key and say what query it makes expensive.
- Name the read and write paths separately; they rarely have the same bottleneck.
- Name the failure you are designing for, then the recovery path.
Follow-up
- What would you drop to keep the system up under load?
- How does this behave when that dependency is down for an hour?
Design an on-demand GPU scheduling system (similar to Instant Clusters…
Design an on-demand GPU scheduling system (similar to Instant Clusters) that dynamically allocates physical hardware to incoming containerized workloads while minimizing fragmentation.
Approach
- Name the failure you are designing for, then the recovery path.
- State the consistency you need, and where you are willing to be stale.
- Fix the scope first: who calls this, how often, and what they do when it fails.
Follow-up
- How does this behave when that dependency is down for an hour?
- What breaks first when traffic grows ten times?
Write the delivery guarantee and replay API for outbound events
The webhook-delivery service fans events to customer endpoints, tracked in webhook_delivery (subscription_id, tenant_id, event_id, redelivery_seq, status, attempt_count, next_attempt_at, lease_token, last_response_code, payload_digest). Customers are asking for exactly-once and in-order delivery. Write the contract you can actually honour: the guarantee, the headers that make it usable, which customer response codes are retryable, the retry schedule and terminal condition, what happens to an endpoint that has been down for a day, and the shape of the redelivery endpoint. State plainly what you are not promising and what the customer must do instead.
Approach
- Refuse exactly-once with the mechanism rather than with policy: the customer's acknowledgement can be lost after they have already committed, so the sender cannot distinguish unprocessed from processed-but-unacknowledged and must retry. The deliverable is at-least-once with a stable event identifier and the customer documented as the deduplicating party.
- Price ordering instead of promising it. With parallel attempts per subscription, a retried event overtakes a newer one, so in-order delivery requires a single in-flight attempt per subscription, which turns any slow endpoint into head-of-line blocking for that subscription's whole queue. Offer it per subscription with that cost written down.
- Define the response contract from the customer's side: 2xx is accepted, 408, 429 and 5xx are retryable, 410 disables the subscription, and every other 4xx is permanent and terminal. Publish the read timeout and tell the customer to acknowledge first and process asynchronously, since their processing time otherwise consumes your worker occupancy.
- Publish the schedule and its end: capped exponential backoff with full jitter, sleeping uniformly in [0, min(cap, base * 2^attempt)] so a mass failure does not re-synchronise the herd, terminal after a stated attempt count or age, plus a per-endpoint circuit breaker that records further deliveries as dropped_circuit_open and notifies the owner rather than burning shared worker capacity.
- Shape redelivery as an insert, not a reset: POST to a redeliveries collection with event ids or a time window creates rows at redelivery_seq + 1 carrying the same bytes, which payload_digest lets you prove, leaving the original terminal rows intact as the record of what happened.
- Compare the event's tenant against the subscription's tenant at enqueue and again immediately before signing, because cross-tenant delivery originates in an enqueue path that took the subscription from one lookup and the payload from another, not in the worker.
Worked solution 30 min
- Write the guarantee sentence and the one-sentence reason exactly-once is not available over HTTP.
- List the headers a customer needs to build their own dedup table: event id, delivery id, attempt number, subscription id, signature and timestamp.
- Write the response-code table with four classes and the action for each, including the 410 case.
- Write the retry schedule with concrete numbers, the terminal condition, and the breaker's threshold and its effect on the delivery row's status.
- Specify the redelivery request and response, and say which columns change and which do not.
Follow-up
- A customer wants everything they missed during their six-hour outage. Do the dropped_circuit_open rows let you answer that, and what retention bound does the answer depend on?
- You offer the ordered mode and one customer's endpoint slows to two seconds per request. What do their delivery metrics look like, and what do you owe them in the docs?
One tenant's counter writes stall the whole connection pool
A change that made a per-tenant usage counter correct now produces site-wide latency whenever one large tenant writes: unrelated endpoints time out waiting for a connection while database CPU stays low and no statement is slow. The change wraps the counter update in a transaction that takes SELECT ... FOR UPDATE on one row, calls an external pricing service, then updates and commits. Give an ordered checklist, the arithmetic that bounds that tenant's write rate, and three repairs with the cost each one accepts.
Approach
- Separate waiting from working. Low database CPU alongside high application latency points at a queue, so instrument connection-acquisition wait separately from query execution time; that queue forms in the application and is invisible in database metrics, which is why the database looks healthy throughout.
- Confirm the lock rather than assuming it: sample waiting sessions and group by wait event, relation and tuple. Contention concentrated on one tuple belonging to one tenant is the signature; a deadlock would instead show the database aborting transactions after its detection timeout, which is not happening here.
- Do the arithmetic out loud. Throughput on a serialised row is one divided by the lock hold time, and the hold spans the external call, so a 20 ms pricing call caps that tenant near 50 writes per second no matter how many pods run. Every waiter also holds a pooled connection while it queues, so the shared pool drains and unrelated tenants fail at acquisition.
- Repair one: shrink the critical section to a single statement with the price resolved before the transaction opens. Cost is a stale price for the duration of one request and a second round trip; benefit is a hold time measured in the database's own execution time.
- Repairs two and three change where the contention lives rather than how long it is held. Sharding the counter into per-(tenant, bucket) rows and summing on read multiplies write throughput by the shard count, at the cost of an aggregate on every read and a shard count you must size against the largest tenant rather than the median. Accumulating in memory and flushing periodically removes the per-write round trip entirely, paid for with a bounded loss window on crash, which is acceptable for a rate limiter and not for a billing counter.
- Contain independently of which repair wins: a separate pool or per-tenant concurrency cap for this write class, a statement timeout low enough that a pathological query dies before it accumulates waiters, and an idle-in-transaction timeout so a stuck client cannot pin a connection and its locks.
Follow-up
- What would a genuine deadlock look like here, which two code paths would produce one, and how does the database's response differ from what you observed?
- If a transaction-pooling proxy sits in front of the database, which of your three repairs changes behaviour, and what stops working that would have worked on a direct connection?
- The counter also enforces a quota. Why is SELECT the count and then INSERT still wrong after you have fixed the contention?
Four days sample coding, design, fundamentals and the practical rounds at deliberately shallow depth, which is enough to surface the topics you did not know were in scope. That map, rather than a guess made on day one, decides where the last three days go.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Map the role and prepare the recruiter screen
- List the must-haves and nice-to-haves the role names (distributed systems, Go, compute/networking/storage fundamentals, PostgreSQL, Kubernetes; QEMU/KVM, KubeVirt, SR-IOV, Infiniband, RDMA, DPUs, CUDA/NCCL, Kafka) and mark each one strong, working or absent for you.
- Write two project summaries for the recruiter screen: one distributed system you ran in production and one infrastructure or low-level piece of work, each with the problem, your change and a measured result.
- Settle a compensation range and write the questions you will ask: which team, which languages are accepted, and which focus the phone screen uses.
- Read every reported question in this guide and rate it: can answer now, can answer after a day's work, or unfamiliar.
Deliverable: A skills map against the role's listed requirements, two project summaries, and the reported question list rated by readiness.
Practice prompt ↗Practice prompt ↗Worked solution ↗02Go concurrency for the technical phone screens
- Build a bounded worker pool in Go with context cancellation, error propagation and clean shutdown; run it under go test -race and fix anything the race detector reports.
- Solve the bank coding titles: detect and break cycles in pod dependencies (three-colour DFS or Kahn's algorithm), schedule GPU pods and drain a node, and split a chunked text stream into line-balanced parts. State the complexity aloud for each.
- Warm up with the first unique character index problem: one pass to count characters, a second pass to find the first with count one, O(n) time.
- For each solution, list the unit tests you would add, including an empty input and, where relevant, a concurrent case.
Deliverable: A race-clean Go worker pool, three bank-style solutions with complexity stated, and a test list for each.
Practice prompt ↗Practice prompt ↗03Systems programming and performance debugging
- Answer aloud: GC vs non-GC memory management for high-throughput network services (allocation rate, GC CPU and pause cost, buffer reuse with sync.Pool, manual lifetime bugs), and file I/O bottlenecks with zero-copy and asynchronous system calls.
- Write a comparison of VFIO, SR-IOV and standard PCIe passthrough: what each shares, what each isolates, and where overhead appears.
- Write an ordered diagnosis checklist for each bank troubleshooting title: a CPU-bound Linux server, storage I/O saturation, slow network throughput including containers, and a server saturated by logging. Name the tool and the metric at each step.
- Work the PostgreSQL throughput-drop question and this guide's connection-pool lock-contention drill: find blocked sessions through pg_stat_activity wait events and pg_blocking_pids(), and separate waiting time from execution time.
Deliverable: Three written systems explanations and four ordered troubleshooting checklists, each naming tool, metric and next step.
Practice prompt ↗Practice prompt ↗04Distributed design: delivery semantics and consensus
- Design the at-least-once task scheduler under network partitions: a durable queue, leases with expiry, fencing tokens so a stale worker cannot commit, and idempotent execution keyed by task id.
- Work the worked exercise 'Write the delivery guarantee and replay API for outbound events' as practice in stating what at-least-once does and does not promise.
- Answer the state synchronisation and consensus question: which state needs consensus (for example etcd for leadership and cluster metadata), which can be eventually consistent, and where regional replicas may serve stale reads.
- End each design with a failure walkthrough: leader lost, partition healed, worker returning after its lease expired.
Deliverable: Two designs, each with its delivery guarantee, consistency choices and a written failure walkthrough.
Practice prompt ↗Practice prompt ↗Worked solution ↗05GPU infrastructure design
- Design on-demand GPU scheduling for containerized workloads: gang allocation for multi-GPU jobs, topology-aware placement, a fragmentation metric, and preemption for high-priority work.
- Compare it with the bank's GPU-aware pod scheduler design and with the reported scenario of short, high-priority inference sharing a GPU pool with long, low-priority fine-tuning jobs.
- Design the global management plane across dozens of data centers: what is global, what stays regional, and how each data center keeps operating while cut off from the global plane.
- Sketch the multi-exabyte object storage design for parallel reads of training data, plus the reported checkpoint-coordination scenario for a 512-node training run.
Deliverable: Three infrastructure designs, each with a placement or data-layout decision, a named failure mode and its recovery path.
Practice prompt ↗Practice prompt ↗06Networking, infrastructure as code and observability
- Write the VLAN vs VXLAN vs VPC comparison: network layer, identifier space (12-bit VLAN IDs vs 24-bit VNIs), overlay vs underlay, and where each fits in multi-tenant isolation.
- Design automated Infiniband partitioning and parallel storage provisioning for bare metal: partition keys configured through the subnet manager, provisioning driven from an inventory source of truth, and verification before a node goes to a tenant.
- Lay out a Terraform and Ansible delivery pipeline: remote state with locking, plan review before apply, staged rollout, drift detection and rollback.
- Design a Prometheus and Grafana node-health pipeline that cordons and drains unhealthy nodes, with rate limits so the automation can never drain a whole cluster.
Deliverable: Written networking comparisons and two pipeline designs, each with the safety check that stops the automation from causing an outage.
Practice prompt ↗Practice prompt ↗07Behavioural stories and a mixed mock loop
- Prepare stories for the reported prompts: a high-severity outage (root cause and prevention), a strong technical disagreement, a design document written while requirements shifted, and replacing urgent manual operations with automation.
- Write a fact sheet per project so scale, team size and timeline stay identical wherever the project comes up.
- Run a mock back to back: one design question from days 4-5, one Go concurrency problem from day 2 and one behavioural story.
- Note where an answer lost structure and redo that segment once.
Deliverable: Four rehearsed behavioural stories with fact sheets, and notes from one mixed mock loop.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
The reported behavioural prompts cover operational ownership, disagreement and ambiguity. Pick stories from infrastructure work where you can name the root cause, the fix and the prevention you added afterwards. For a disagreement, show the evidence you used and how you moved forward once a decision was made. For a design document under shifting requirements, say what you fixed early, what you left open, and how you kept reviewers current.
Tell me about a situation where you had a strong technical disagreemen…
Tell me about a situation where you had a strong technical disagreement with a peer or lead. How did you present your arguments, and how was the conflict resolved?
Approach
- Pick a story where you made the decision, not one where you watched it.
- Give the blast radius: what could have broken, and what you measured.
- Close with what you would do differently, concretely.
Follow-up
- How did you know your change caused the improvement?
- What would you do differently if you ran that again?
Describe a time you had to resolve a high-severity production outage u…
Describe a time you had to resolve a high-severity production outage under intense time constraints. How did you identify the root cause, and what steps did you take to prevent recurrence?
Approach
- Name the disagreement and how you resolved it with evidence.
- State the situation in two sentences and spend the rest on the reasoning.
- Give the blast radius: what could have broken, and what you measured.
Follow-up
- How did you know your change caused the improvement?
- What did you decide not to do, and why?
Reverse a webhook ordering decision after measuring its cost
You argued for strict per-subscription ordering in webhook-delivery, which means one in-flight attempt per subscription. It shipped. Three months later a single unresponsive endpoint holds one subscription's queue at a six-hour backlog, and two customers report events arriving out of order anyway once their own retries are counted. Describe a decision you reversed: what you originally optimised for, the measurement that changed your mind, what the reversal cost in engineering time and customer change, and how you told the people who had already built on the original guarantee.
Approach
- State the original decision as a trade you made knowingly. Ordering across a network requires a single in-flight attempt per subscription, and its price is head-of-line blocking whenever one endpoint is slow. 'We priced it wrong' is a much stronger opening than 'we did not realise', and it is usually the true one.
- Bring the measurement that flipped it, not the anecdote: backlog age at the ninety-ninth percentile per subscription, the share of subscriptions where one slow endpoint gated an otherwise healthy queue, and the delivery throughput lost to serialisation. A reversal justified by complaints is indistinguishable from a reversal justified by fatigue.
- Name what you learned about the guarantee itself, which is the engineering content of this story. At-least-once delivery means a retried event already arrives after newer ones and the consumer already must be idempotent, so a guarantee the customer has to defend against anyway was never worth what it cost to provide.
- Describe the migration, because reversing a published contract is the hard half and the part candidates skip. Parallel attempts behind a per-subscription flag, a monotonically increasing sequence number added to the envelope so order-sensitive consumers can sort or discard, documentation that states at-least-once and unordered in those words, and a deprecation measured in quarters because the client is a pinned SDK inside a build pipeline you cannot see or redeploy.
- Give the cost in the two currencies that matter: engineer-weeks, and how many customers had to change code. Then say who you told before it shipped rather than in a changelog afterwards, and which large customer you left on the old behaviour and for how long.
- Close with the signal you now weight differently, stated as something you would do earlier next time: measuring the blocking cost on the slowest decile of endpoints before committing to the guarantee, rather than after a customer noticed.
Follow-up
- A customer insists they need ordering. What do you offer them that is not global serialisation?
- How did you choose the deprecation window given that you cannot see or redeploy the clients?
- What would have to be true for you to reverse back?
- 01
Describe a time you had to resolve a high-severity production outage under intense time constraints. How did you identify the root cause, and what steps did you take to prevent recurrence?
- 02
Tell me about a situation where you had a strong technical disagreement with a peer or lead. How did you present your arguments, and how was the conflict resolved?
- 03
How do you approach writing a technical design document for a highly complex, cross-functional project where requirements are still shifting?
- 04
Tell me about a time you replaced urgent manual operations with automation while keeping the system stable and delivery on track.
Is this an official Together Ai interview guide?
No. It is PracHub's own research and practice material for the Software Engineer role at Together Ai. Rounds and questions reflect what candidates have reported, not a process Together Ai has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗Which programming language should I use during the technical interviews?
Reports say Together AI's backend and cloud infrastructure work is primarily in Go, and that any modern backend language such as Go, C++, Rust or Python is generally accepted in coding assessments. Comfort with Go still helps: Go-specific questions on concurrency and performance are reported in systems-focused rounds, and one reported question asks how you would build a custom Kubernetes operator in Go. Confirm the accepted languages with your recruiter.
PracHub interview research ↗How deep does the systems design discussion go?
Reportedly deep: candidates are expected to connect the design to hardware and network resources, including how memory is managed, how packets are routed and how database transactions are isolated. Prepare by taking each design at least one layer below the box diagram: how the scheduler learns GPU topology, how storage reads are parallelised, and how tenant isolation is enforced in the network.
PracHub interview research ↗What is the typical timeline for the interview process?
Candidates report roughly 3 to 5 weeks from recruiter screen to final decision, across three stages: a recruiter screen, one or more technical phone screens, and a virtual or on-site loop. Scheduling and team can change this. If you have heard nothing after the final loop, follow up with your recruiter.
PracHub interview research ↗Is there algorithm-style coding as well as systems questions?
Yes. The question bank for this role includes coding titles such as detecting and breaking cycles in pod dependencies, scheduling GPU pods and draining a node, splitting a chunked text stream into line-balanced parts, and returning the index of the first unique character. Practise graph traversal, hashing and careful simulation, write the solutions in the language you plan to use, and state their complexity.
PracHub Software Engineer practice ↗Do I need GPU virtualization or Infiniband experience?
The role lists them as nice-to-haves, alongside QEMU/KVM, KubeVirt, SR-IOV, RDMA, DPUs and CUDA/NCCL. The must-haves are distributed systems, a backend language, systems fundamentals, PostgreSQL and Kubernetes. Reported questions still touch VFIO vs SR-IOV vs PCIe passthrough and Infiniband partitioning, so learn the concepts well enough to reason about the trade-offs even without hands-on experience.
PracHub Software Engineer practice ↗What kind of troubleshooting questions should I expect?
Reported and bank questions include debugging a PostgreSQL throughput drop under concurrent writes, diagnosing CPU problems on a Linux server, diagnosing storage I/O and identifying the workload, slow network throughput including containers, and a server saturated by logging. Practise an ordered method for each: the resource, the tool, the metric, and what reading rules a cause out.
PracHub Software Engineer practice ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-24 - 02PracHub Software Engineer practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-24 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-24