Reliable LLM tool calling: format control, failure recovery, and runaway loops
Company: UiPath
Role: Machine Learning Engineer
Category: Machine Learning
Difficulty: medium
Interview Round: Technical Screen
You are building an LLM agent that uses tools: functions or APIs the model can invoke while working on a task. Answer the following questions about tool calling in order.
### Clarifying Questions
- Is the model a hosted API with built-in function calling, or a self-hosted open-weights model whose decoding you control?
- Do the tools only read data, or do some have side effects such as writing records or sending messages?
- What does "too long" mean for a tool-call loop here: too many steps, too many tokens of context, too much latency, or too much cost?
- How many tools does the agent have access to, and do they change often?
### Part 1 — How an LLM calls tools
Explain end to end how a language model calls a tool: what the model sees, what it produces, who actually executes the tool, and how the result gets back to the model.
```hint Who runs the code
The model only produces tokens; think about which component turns those tokens into an actual function call.
```
#### What This Part Should Cover
- Tool schemas and descriptions in the context, and the model emitting a structured call
- The runtime loop: parse, validate, execute, append the result, call the model again
- How models learn this behavior, and single versus parallel calls
### Part 2 — Controlling the tool-call format
How do you make the model's tool calls follow the required format, such as valid JSON that matches each tool's argument schema?
```hint Before, during, and after
Consider what you can do in the prompt, what you can do while tokens are being generated, and what you can do once the output exists.
```
#### What This Part Should Cover
- Schema design and prompting: clear names, types, enums, examples
- Constrained or structured decoding and native function-calling modes
- Validation and repair of malformed output
### Part 3 — Stability and fixing incorrect tool use
How do you make tool calling stable across runs, and what do you do when the model fails to call tools correctly (the wrong tool, wrong arguments, or no call when one is needed)?
```hint Diagnose before fixing
Different failure types call for different fixes; start by measuring which ones actually occur.
```
#### What This Part Should Cover
- Evaluation sets and metrics for tool use, and categorizing failures
- Inference-time fixes: retries with error feedback, fewer or better-described tools, tool routing, decoding settings
- Training-time fixes such as fine-tuning on corrected traces, and guarding side-effecting tools
### Part 4 — A tool-call loop that runs too long
How do you handle an agent whose tool-call loop goes on for too long?
```hint Why it keeps going
Separate loops caused by repetition, by missing information, and by an overly broad task before choosing a control.
```
#### What This Part Should Cover
- Hard limits (steps, tokens, time, cost) and what happens when one is reached
- Detecting unproductive loops such as repeated identical calls
- Keeping context manageable in long loops and ending gracefully
### What a Strong Answer Covers
- A correct mental model: the model proposes calls and the application executes them
- Layered format control (prompt, decoding, validation) rather than a single trick
- Measurement-driven reliability work with both inference-time and training-time remedies
- Bounded, observable agent loops that degrade gracefully and protect side-effecting actions
### Follow-up Questions
- How would you build an evaluation set for tool calling, and which metrics would you report?
- A tool returns a very large result that crowds the context window. What do you do?
- How do you prevent a retried tool call from executing a side effect twice?
- When would you fine-tune the model instead of improving prompts and tool descriptions?
Overview: An LLM agent question on tool calling: how a model invokes tools, how to keep tool calls in a valid format, how to make tool use stable and fix wrong calls, and how to handle tool-call loops that run too long. It tests practical reliability engineering for production agents.
Read the full UiPath Machine Learning Engineer interview experience this question came from