AI & agents
Every AI system is still a software system
The interfaces, data, permissions and tests that turn a model into a useful product feature.
An AI feature becomes useful through the software around it. The model may interpret a request, extract information or draft an answer. The application decides what data it receives, what actions it can request, and how a person checks and uses the result.
Those decisions matter even when a demonstration looks convincing. A good answer from the wrong source can mislead. A correct proposed change can still be applied to the wrong record. A useful result can arrive too late or require more checking than the task previously took.
Start with the work the feature should improve, then design the model's role within it.
Define the useful result
Name the user, the task and the next action. “Summarise documents” leaves important questions open. A summary for someone deciding whether to read a report needs different detail from one used to prepare a record change.
Describe what success would let the person do. Establish what information must be present, which errors matter, and when the system should ask for help. Include the time spent reviewing and correcting the output in your comparison with the existing process.
Consider the interface at the same time. The result might belong beside a field, in a comparison table, or in an editable document. Conversation can help people express an unfamiliar request, but it need not become the destination for every result.
UX is part of the system offers a way to choose where a capability belongs. That choice can reduce the complexity of the AI feature before any model is selected.
Keep exact work in software
Use ordinary code for calculations, stable rules, validation and allowed state changes. Models can help interpret varied input or propose a route through a task; they do not need to reproduce a calculation that the application can perform exactly.
Consider a constructed example in which an assistant prepares a record update. The model might interpret the user's intention and propose field values. The application still checks the record identifier, field types, current version and user's permissions before accepting the change.
Those checks should apply regardless of how confidently the model explains its proposal. A prompt can describe a rule, but the tool or application must enforce it.
This division also helps testing. A developer can test validation and access controls independently of the model, then evaluate whether the model supplies suitable proposals. A failure in one layer should not be hidden by a plausible result from another.
Make source data usable
For organisation-specific answers, the system needs reliable access to relevant records. A model's learned knowledge does not establish the current state of an organisation or the authority of a supplied document.
Check what the sources mean, how current they are, and which source should prevail when they disagree. Decide what the product should do when information is missing. Preserve source references where a person needs to verify a claim.
Retrieval helps find material; it does not make that material correct. A search result may contain an obsolete procedure, an incomplete record or instructions that should be treated as untrusted content. The application needs rules for selecting and using it.
The data path also includes caches, logs and evaluation records. Keep only what serves a defined purpose, with suitable access and retention. Debugging an AI feature should not create a less protected copy of its source data.
Enforce permissions before reading and acting
An assistant should not reveal a record that the user cannot access through the application. Check access before assembling model context. Check action permissions again before changing state.
Reading and writing are different authorities. A person who can view an order may not be allowed to approve or delete it. Tools should expose the particular operations the feature needs, with validation and understandable failure states.
Plan for integration failures too. Credentials expire, requests time out, and a sequence can stop after only some changes have completed. Define how retries avoid duplicate effects and how users find out what actually happened.
Provider selection follows these requirements. Data location, retention, access and operating cost can exclude an otherwise capable option. Choosing an AI provider around your data requirements separates provider commitments from controls the application must implement.
Evaluate the complete task
Compare the feature with a credible baseline on representative cases. Test routine work, ambiguous input, missing evidence, permissions and failures in connected systems. Include cases where asking a question or declining an action is correct.
Measure the result the user needs, not only the quality of generated text. An assistant may select a good source but misread it. It may propose the correct change with the wrong identifier. A reviewer may struggle to spot either error if the interface hides the supporting evidence.
For agents, inspect the meaningful steps as well as the final answer. When an agent earns its place explains when adaptive tool selection is useful. Evaluating an AI feature before release describes how to turn that assessment into a release decision.
Give the feature an owner after release
Real use brings new inputs and reveals gaps in the original evaluation. Dependencies, models and source data also change. Define who investigates failures, maintains the tests and approves changes.
Monitor signals that relate to the task: missing sources, failed tool calls, corrections, deferrals, latency and cost. A successful HTTP response cannot tell you whether the person received a useful answer. Equally, a user edit does not always mean the model failed; the requirement or available information may have changed.
Keep the operating arrangement proportionate. A small read-only feature and a system that changes important records need different controls and support. Both need an explicit account of who maintains them.
The NIST AI Risk Management Framework provides a broader voluntary framework for managing AI risks through its lifecycle. Applying the practices in this article is an engineering approach, not a claim of conformity. The immediate question remains practical: can this complete feature improve the task under the conditions in which people will use it?