AI & agents
When an agent earns its place
Compare adaptive tool selection with simpler designs, then define tools, review and stopping.
An agent chooses part of its route while a task is underway. It might inspect a source, choose a tool, examine the result and decide what to do next. That flexibility is useful when the right next step depends on information that is not available at the start.
A workflow can also branch, use models and handle varied inputs. The distinction here is who controls the sequence: rules written by the development team, or a model choosing some of the steps at runtime.
Choose an agent when that adaptive choice improves the task enough to justify the additional uncertainty, latency and maintenance. Begin with the simplest design that can do the work well.
Identify the choice you are delegating
Describe what the agent would decide. It may choose which source to search, which diagnostic check to run, or whether it has enough evidence to prepare a draft. “Help with operations” does not identify a decision that can be evaluated.
Specify the outcome, available information and permitted tools. State what would make the task complete, when more information is needed, and which actions remain with a person.
A useful candidate has a clear outcome even when the route varies. Investigating an unfamiliar error may require different checks depending on the first result. Moving a validated record through a fixed approval sequence usually needs an ordinary workflow.
The whole product need not be agentic. A small adaptive loop can operate inside an application that handles identity, validation, storage and state changes through conventional software.
Compare the alternatives on the same task
Compare a conventional interface or workflow, AI assistance within a defined sequence, and an agent that selects tools. Use the same inputs and completion criteria.
An agent may avoid a complicated set of hand-written branches. It may also make unnecessary calls or take a different route on repeated attempts. Measure time to a usable result, errors, review effort and cost across the whole task.
Varied language alone does not require an agent. A model can classify a request or extract fields before a fixed workflow proceeds. Tool selection becomes useful when the intermediate results genuinely affect what should happen next.
Feedback must be available. If an action returns no meaningful result, the agent has little basis for deciding whether to continue. Break long sequences into observable steps, and make missing or uncertain results explicit.
Give tools narrow responsibilities
Tools should express allowed operations rather than unrestricted access to a system. Prefer “prepare a record change” to “edit the database”. Validate the request at the tool boundary even if the model has already checked it.
The following is an illustrative contract, not a description of a client system:
| Part | Responsibility |
|---|---|
| Input | Record identifier, expected version, permitted field changes and reason |
| Checks | User access, field types, allowed transitions and whether the record has changed |
| Result | A proposal for review, a validation error or a version conflict |
| Commit | A separate operation checks approval and the current record before writing |
Approval should cover the action actually executed. If the target or proposed values change after review, check whether new approval is required. Do not let a model turn permission for one change into permission for another.
Retries need deliberate handling. An operation should use an appropriate idempotency mechanism when repeating it could create duplicate effects. A timeout may mean the action failed, or that it succeeded but the response was lost. The tool must help the application resolve that uncertainty.
Evaluate actions and stopping
A polished final answer can conceal a poor sequence. Inspect relevant source selection, tool arguments, permission checks, retries and the stopping reason.
Build cases in which the correct result is to ask for information, defer to a person or decline an action. Score those outcomes against the case requirements. Completion rate alone cannot tell you whether the system respected its boundaries.
Give the agent explicit limits on attempts, cost and elapsed work. It should also stop when a required permission is absent, sources conflict beyond its ability to resolve them, or a tool leaves state uncertain. The final report should distinguish a prepared action from a completed one.
Test untrusted content that tries to redirect the task. NIST's agent-hijacking evaluation work shows why known attacks alone are insufficient and why repeated attempts and task-specific results matter. A passing test set supports a limited claim about the tested system; it does not prove that prompt injection is solved.
Evaluating an AI feature before release provides the broader method for constructing cases and deciding what the results support.
Make review useful
If a person must approve an action, show what will change, which records are affected and what evidence supports it. Give them enough context to check the important difference without repeating the entire investigation.
Review can become superficial when proposals are frequent and plausible. Treat that as a risk to measure: can reviewers detect the errors that matter, and how long does checking take? Stable limits and malformed inputs should be checked in software rather than left for a person to notice.
Start with the least authority that provides value. Read-only investigation or proposed changes may be useful releases in their own right. Greater authority should follow evidence about the current task and its failures, with a way to reduce it again when conditions change.
An agent earns its place through the work it improves. Keep the delegated choice small enough to explain, the tools constrained enough to test, and the outcome clear enough that a person can tell what happened.