Every rule this product enforces came out of a run that went wrong. These are the notes that argue for them — written up as we hit them, most of them clickable rather than described. Read them as the working, not the marketing: if you disagree with one, you have found the place to push on the product.
Agent TrustKit asks you to accept a gate decision about your own agent. That is only worth something if you can see the reasoning it was built on.
The product's whole argument is that a score with no evidence behind it is not a claim anyone can sign. The same standard applies to the product itself. Every dimension in a clearance report encodes a position — that a green trace is not a passing run, that a correct answer from contradicted evidence is a failure, that a frozen answer key books your improvements as regressions. Those are arguable positions, and you should be able to read the argument before you trust the gate.
So they are published in full, with the runs they came from. They are also the honest version of the sales pitch: a note that shows a config winning on silhouette by discarding 85% of the corpus is not flattering to anyone, including us.
A note on where they live. Agent TrustKit is a product of O&A Consulting, and the field notes are published on the studio's site, where they have always lived. Each one below opens on oandaconsult.com in a new tab — same authors, same work, different masthead.
Expected: yes. Actual: yes. PASS. Then you open the trace and the evidence says the opposite. Watch a passing test come apart — then the harder problem, because the trace is the agent's own testimony.
Why groundedness is scored against what the agent actually retrieved on that run, not against a reference answer — and why the report treats the trace as testimony rather than proof.
A graph and a Cypher prompt turns retrieval into a guessing game. Watch the same question answered twice — five queries and four dead ends on one side, two tool calls on the other — then pull out the generic query tool and see what breaks.
Why tool use is scored as four separate questions — allowed, right tool, right arguments, result used correctly — instead of one number. The failure in this note is invisible to any of the other three. The evaluation guide has the full set.
Model strength pays off in proportion to how well-posed your question is, and below a threshold it stops paying at all. Pick the step you'd blame in a failed agent run, then watch a judge upgrade cash out on one question and evaporate on another.
The argument for the task contract. If the question isn't determinate before it is asked, a stronger grader buys you nothing — which is why the model here extracts evidence and never computes a score. How it works.
Chunking decides what your clusters are about. Cleanup decides whether they're about anything. Measured on 31,000 verses, including the config that won on silhouette by throwing away 85% of the corpus.
Why a groundedness score is only as good as what reached the agent, and why a retrieval finding is reported as a retrieval finding rather than charged to the agent's reasoning.
Where golden datasets earn their keep in agentic evaluation, where they fall apart, and what we use instead. Two slides you can click: watch a good change turn a test red, then catch an answer key that's wrong.
Why a regression is a diff against the contract's dimensions rather than against a frozen expected answer — the mechanism behind continuous evaluation and trustkit compare.
The notes argue. These answer — each one states a position in the first paragraph and then points at the note that makes the case properly. Also on oandaconsult.com, where the full set lives.
The notes are the reasoning. The rest of this site is what it turned into: how the evaluation runs and why the model never computes a score, the evaluation guide for the six dimensions, agent clearance for what a signed scope actually says, and four sample reports on the same agent — including the two that don't pass.
Send one trace export and a sentence on what the agent was supposed to do. You get the report back, free, whatever it says.
Book a call →[email protected]