Evaluation
Evaluating an AI feature before release
Define the task, compare a baseline and test the failures that matter to the release decision.
Evaluation should help a team decide whether an AI feature is ready for a particular use. Start with the task, the errors that matter and the conditions under which the feature will operate. Then test the complete application against those requirements.
A model benchmark can inform selection. It cannot establish that your retrieval, permissions, tools and interface work together for your users. A convincing demonstration is useful evidence of possibility, but it is a poor estimate of reliability across unfamiliar cases.
The release decision needs a defined scope and enough evidence to judge it. The following method can scale from a small assisted feature to an agent that uses several tools.
Define success before tuning
Write a short account of what the feature may do. Name its users, inputs, outputs and permitted actions. State what constitutes a material error and when asking for clarification or declining to act is correct.
For an illustrative record-update assistant, success might mean selecting the right record, proposing permitted fields, explaining the source of each change and leaving the write for approval. Producing a plausible description of the update would not be enough.
Use observable criteria. “Helpful and trustworthy” is difficult to score consistently. “The proposal identifies the record, preserves unchanged fields and defers when the source conflicts” gives reviewers specific behaviour to inspect.
Decide which criteria block release and which can be improved gradually. An access-control failure should not disappear inside an average score that rewards fluent language or faster completion.
Compare with the existing work
Choose a credible baseline: the current manual process, an existing product version, or a simpler interface with deterministic rules. Compare the same task from preparation through review and correction.
An AI feature should not receive credit for the convenient part of a task while the baseline includes the work around it. If users must gather sources before prompting or repair the result afterwards, include that effort.
Useful measures depend on the task. They may include correct completion, material error rates, appropriate deferral, time to a usable result, review effort and operating cost. Keep them separate when combining them would conceal a significant trade-off.
A slower feature may still be worthwhile if it improves accuracy or access. A faster one may be unsuitable if mistakes become harder to detect. The evaluation should make that choice visible rather than deciding it through an unexplained score.
Build cases that represent the supported scope
List the task types and conditions the feature must handle. Include common inputs, legitimate variation, missing information, conflicting sources and cases outside scope. Add permission boundaries, tool errors and partial completion where those can affect the outcome.
Record why each case exists. It may come from an observed failure, a requirement or a constructed scenario. Synthetic cases can test a mechanism, but they do not establish how frequently it occurs in use.
Separate cases used during development from cases held back for release review. Once developers inspect a held-out failure and tune against it, that case is useful regression material rather than fresh evidence of generalisation. Keep a new assessment set when that distinction controls the decision.
For variable model behaviour, repeat consequential cases and record the distribution of outcomes. One successful run can hide intermittent failure. Repeated runs of one case also do not substitute for coverage across different tasks and users.
Keep test data governed. Use approved records, suitable redaction or constructed inputs that preserve the relevant difficulty. Evaluation should not create an unmanaged second copy of production data.
Test the layers and the whole path
Test calculations, parsers, permissions, validation and state transitions as ordinary software. These checks help locate defects without attributing every failure to the model.
Next test the model's bounded responsibility: extraction, classification, evidence selection or proposed tool arguments. Assess properties that matter to the task rather than requiring one exact wording when several answers are valid.
Run complete scenarios through the application's context, retrieval, interface and connected services. Check both what the user sees and what state actually changes. A success message can be wrong even when the underlying tool behaved correctly.
For agents, inspect significant steps: sources accessed, tools selected, arguments, approvals, retries and stopping. Include malicious instructions in untrusted content and attempts to cross user or tenant boundaries. Controls must be enforced by the surrounding software.
NIST's agent-hijacking evaluation work highlights the value of adaptive attacks, task-specific analysis and multiple attempts. Use those findings to challenge the tested system rather than treating a fixed list of attacks as complete.
Make human judgement consistent
Give reviewers criteria before they see outputs. Include examples of acceptable variation, missing evidence and errors that must block acceptance. Have reviewers assess some common cases and investigate disagreements.
A disagreement may reveal an unclear rubric or a task with more than one reasonable answer. It may also expose missing domain knowledge. Do not resolve every disagreement by averaging scores and moving on.
If a model assists with grading, check its judgements against suitable human review and known cases. An automated judge can have its own biases and failure modes, particularly when assessing plausible but unsupported answers.
Measure review effort too. If a reviewer must reconstruct every source and intermediate step, the feature may be moving work rather than reducing it. Improve the evidence presented before assuming users can absorb that cost.
Decide what the result supports
Analyse failures by both mechanism and consequence. Missing source data, incorrect retrieval, a model error and a broken permission check require different repairs. Prioritise the failures that matter within the intended scope.
Record the application version, model configuration, test-set version, conditions and results. A finite sample leaves uncertainty, especially for rare failures. Report counts and coverage alongside percentages; avoid presenting a small sample as a precise population estimate.
“No prohibited failures observed” is a statement about those tests. It does not prove zero risk. A release policy may require that result while also limiting users, data or action authority until more operating evidence exists.
Check recovery before release. Exercise rollback or a reduced-function path, and establish who responds to failures. A staged release can begin with read-only assistance or proposed actions while preserving a useful service.
Continue learning after release
Monitor known failure modes and investigate corrections, deferrals and changes in task performance. Preserve enough context to distinguish an error from a changed user requirement.
Add confirmed failure patterns to regression tests. Reassess relevant cases when models, prompts, retrieval or tools change. Keep the supported scope accurate as people find new uses.
The NIST AI Risk Management Framework treats risk management as work across the lifecycle. This article's method supports a particular engineering decision; it does not certify a product or establish permanent reliability. Release the scope the evidence supports, then gather evidence before expanding it.