Findings are what Agent TrustKit's checks could verify against the contract and the evidence available in the trace, not an exhaustive audit. A PASSED gate means nothing checked failed; it does not certify that no other issue exists.
| Dimension | Violation rate | Threshold | Violations | Warnings | Status |
|---|---|---|---|---|---|
| task completion | 2.5% (1/40 runs) | ≤ 5% | 1 | 0 | pass |
| groundedness | unevaluated (0/40 skipped) | ≤ 0% | 0 | 0 | unevaluated |
| completeness | unevaluated (0/40 skipped) | ≤ 0% | 0 | 0 | unevaluated |
| constraint adherence | 0.0% (0/40 runs) | ≤ 0% | 0 | 0 | pass |
| tool use quality | 0.0% (0/40 runs) | ≤ 5% | 0 | 1 | pass |
| efficiency | 5.0% (2/40 runs) | ≤ 5% | 2 | 0 | pass |
| recovery escalation | 2.5% (1/40 runs) | ≤ 5% | 1 | 0 | pass |
| Metric | Mean | Median | P95 | Max |
|---|---|---|---|---|
| Cost per run (USD) | $0.0374 | $0.0221 | $0.0336 | $0.6100 |
| Latency per run (ms) | 4952 | 3567 | 4654 | 61200 |
Per-tool execution reliability across the runs: how often each tool was called and how often the call errored. Deterministic from tool-call error signals, over every observed tool. Informational. Never gates.
Tool calls: 49 · failures: 4 (8.2%)
| Tool | Invocations | Failures | Failure rate |
|---|---|---|---|
| get_document | 45 | 4 | 8.9% |
| web_lookup | 4 | 0 | 0.0% |
| Task | Trials | Trials w/ violations | Tool-set similarity | Step-count CV | Cost CV | Duration CV |
|---|---|---|---|---|---|---|
| brief-meta ⚠ mixed | 5 | 1 | 1.00 | 0.00 | 0.23 | 1.57 |
| brief-nvda ⚠ mixed | 5 | 1 | 0.80 | 0.12 | 1.70 | 0.12 |
| brief-tsla ⚠ mixed | 5 | 1 | 1.00 | 0.00 | 0.54 | 0.15 |
| brief-aapl | 5 | 0 | 0.80 | 0.14 | 0.07 | 0.13 |
| brief-amzn | 5 | 0 | 1.00 | 0.00 | 0.16 | 0.23 |
| brief-goog | 5 | 0 | 0.80 | 0.12 | 0.20 | 0.16 |
| brief-jpm | 5 | 0 | 1.00 | 0.24 | 0.10 | 0.19 |
| brief-msft | 5 | 0 | 0.80 | 0.14 | 0.23 | 0.30 |
⚠ 3 of 8 tasks are mixed: some trials violate, some don't. A pooled violation rate hides this; mixed tasks are where the agent is unpredictable.
Informational. Never gates. Tool-set similarity: 1.00 means identical tools every trial. CV = stdev/mean; higher = less repeatable.
Did the agent stay within cost, latency, and step budgets?
nvda-t2 · checker
[email protected]meta-t1 · checker
[email protected]Did the agent achieve the stated objective and produce the required output?
tsla-t3 · checker
[email protected]When steps failed, did the agent recover or escalate rather than spiral?
tsla-t3 · checker
[email protected]e2 · timeout fetching documentDid the agent use the right tools, look in the right places, and avoid redundant calls?
jpm-t0 · checker
[email protected]r0 · first of 3 identical calls, input={'id': 'jpm-note-1'}aapl-t3w1msft-t4w1tsla-t3e1e2tsla-t3e1e2