Evaluation produces a score. Clearance is the decision — the moment a named person accepts that an agent is fit to run against real users, real money, or real records, and puts their name on that judgement. A number does not make that decision. Evidence, organised so it can be checked rather than trusted, is what does.
"The agent scores 94%" is not an answer to "are you comfortable turning this on."
Every agent project reaches the same room. The engineers have metrics. The person who has to authorise the launch has a different question, and it is not a quantitative one: what happens when this is wrong, how would we know, and who is accountable. A score is silent on all three.
So the meeting stalls, and it stalls in a predictable way. The engineering side experiences it as risk aversion — the numbers are good, what more do you want. The risk side experiences it as being handed a number with no way to interrogate it. Both are behaving reasonably. The artifact is wrong for the decision.
What actually unblocks that room is not a better score. It is evidence a reviewer can check in the time they have, mapped to the things they are accountable for, with the failures shown rather than averaged away.
Produces measurements. Scores per dimension, per run, over time. The audience is the team building the agent, and the question it answers is "is this getting better."
Produces a decision and its justification. The audience is whoever carries the consequence, and the question is "can we turn this on, and what did we accept when we did."
It runs forever, on a sample of traffic, and its value is the trend line. Nothing about it terminates.
It happens at a moment, against a stated scope, and it expires — because the model, the corpus, and the traffic all move underneath it.
Evaluation is a precondition for clearance and not a substitute for it. You cannot clear an agent you have not evaluated. You can very easily evaluate an agent thoroughly and still have produced nothing anyone will sign, which is the situation most teams are actually in.
What the agent is cleared to do, and against what population. An agent cleared for internal research is not cleared for client-facing answers. Scope creep after clearance is the most common way a signed decision quietly stops being true.
What success meant, written down before the run — a contract. Criteria assembled after seeing the results are not criteria, and a reviewer has no way to tell the difference unless the document is dated and versioned.
Every claim in the package resolves to the span of the trace it came from. A reviewer should be able to spot-check any single line in under a minute. Anything that can only be believed, rather than checked, is doing no work.
Which runs failed, how, and what it would take to fix each one. A package that shows only aggregate success is asking to be trusted. The failures are the most informative thing in it, and hiding them is what makes reviewers distrust the rest.
Where the evidence runs out: thin telemetry, untested paths, populations not represented. Stating this is what makes the rest of the document credible, and it is the section a good reviewer reads first.
Who signed, when, and when it must be revisited. A clearance with no expiry becomes a claim about a system that no longer exists — new model version, new corpus, new traffic.
Clearance is not automatable, and the instinct to automate it is the failure mode. If a system could decide on its own that it was safe to deploy, the accountability would have nowhere to sit — and accountability with nowhere to sit is the thing every governance framework exists to prevent.
What is automatable is everything that makes the decision cheap: gathering the evidence, running the deterministic checks, scoring what needs scoring, organising it against the criteria, and surfacing the failures first. Done well, that turns a week of a senior engineer's time into an hour of a reviewer's — without moving the judgement itself off a human.
This is also why the reviewer is usually not the builder. The person who wrote the agent is the worst-placed person to certify it, for the same reason authors do not review their own pull requests.
The framework's Measure and Manage functions ask for documented, repeatable evidence that risks were characterised and addressed, with accountability assigned. A clearance package is that evidence for one agent at one moment — which is why building it as a byproduct of evaluation is far cheaper than reconstructing it later for an audit that arrives without warning.
A package nobody could realistically check, signed because the meeting needed to end. Recognisable by its length and its absence of named failures. It transfers blame without reducing risk.
Signed once, never revisited, while the model provider ships two new versions underneath it. The document still exists; the system it describes does not.
Results first, standards second. Always passes, because the bar was placed where the results already were. Only a dated, versioned contract distinguishes this from the real thing.
Evidence gathered on curated cases, then generalised to production traffic that looks nothing like them. The scope section exists to prevent exactly this, which is why it belongs first.
The expensive version of clearance is the one assembled reactively, when a risk team asks for evidence about an agent that has been running for six months and whose traces have long since rotated out of retention. At that point the work is archaeology.
The cheap version falls out of continuous evaluation: if criteria are already written down, evidence is already attached to scores, and failures are already named, then a clearance package is a view over material you have been producing all along. The decision stays human. The dossier assembles itself.
Agent TrustKit produces exactly this artifact: contract-driven scoring where every judgement quotes its evidence, failures surfaced rather than averaged, limits stated plainly, and the whole thing generated from the OpenTelemetry your agents already emit — on your machine, against your data, offline.
The first report is free. Bring a contract and the telemetry you already have.
Book a call →[email protected]