← All four sample reports

Trust Report — pm-morning-briefing-agent

Agent: atlas-briefing-agent · 40 runs evaluated · generated 2026-08-16 02:30 UTC

Gate: INCONCLUSIVE — unevaluated (checks skipped): groundedness, completeness Dimensions with a contract threshold pass at violation_rate <= threshold; all other dimensions are zero-tolerance for violations. A dimension whose checks were skipped on every run, or for which no checker ran at all, is unevaluated and blocks a clean pass.

Findings are what Agent TrustKit's checks could verify against the contract and the evidence available in the trace, not an exhaustive audit. A PASSED gate means nothing checked failed; it does not certify that no other issue exists.

Trust profile

DimensionViolation rateThreshold ViolationsWarningsStatus
task completion 2.5% (1/40 runs) ≤ 5% 1 0 pass
groundedness unevaluated (0/40 skipped) ≤ 0% 0 0 unevaluated
completeness unevaluated (0/40 skipped) ≤ 0% 0 0 unevaluated
constraint adherence 0.0% (0/40 runs) ≤ 0% 0 0 pass
tool use quality 0.0% (0/40 runs) ≤ 5% 0 1 pass
efficiency 2.5% (1/40 runs) ≤ 5% 1 0 pass
recovery escalation 2.5% (1/40 runs) ≤ 5% 1 0 pass

Efficiency

MetricMeanMedian P95Max
Cost per run (USD)$0.0390 $0.0229 $0.0338 $0.6100
Latency per run (ms)3489 3714 4687 4733

Tool reliability

Per-tool execution reliability across the runs: how often each tool was called and how often the call errored. Deterministic from tool-call error signals, over every observed tool. Informational. Never gates.

Tool calls: 46 · failures: 4 (8.7%)

ToolInvocations FailuresFailure rate
get_document45 4 8.9%
web_lookup1 0 0.0%

Cross-trial consistency

TaskTrials Trials w/ violations Tool-set similarityStep-count CV Cost CVDuration CV
brief-nvda ⚠ mixed 51 0.80 0.12 1.65 0.28
brief-tsla ⚠ mixed 51 1.00 0.00 0.54 0.12
brief-aapl 50 1.00 0.12 0.13 0.29
brief-amzn 50 1.00 0.00 0.12 0.26
brief-goog 50 1.00 0.00 0.23 0.19
brief-jpm 50 1.00 0.24 0.24 0.08
brief-meta 50 1.00 0.00 0.17 0.11
brief-msft 50 1.00 0.12 0.23 0.19

⚠ 2 of 8 tasks are mixed: some trials violate, some don't. A pooled violation rate hides this; mixed tasks are where the agent is unpredictable.

Informational. Never gates. Tool-set similarity: 1.00 means identical tools every trial. CV = stdev/mean; higher = less repeatable.

Scope

Findings

Efficiency

Did the agent stay within cost, latency, and step budgets?

VIOLATIONRun cost $0.6100 exceeded budget $0.5000
run nvda-t2 · checker [email protected]

Task Completion

Did the agent achieve the stated objective and produce the required output?

VIOLATIONRun produced no final output
run tsla-t3 · checker [email protected]

Recovery Escalation

When steps failed, did the agent recover or escalate rather than spiral?

VIOLATIONRun ended in an unrecovered error after 2 failed step(s)
run tsla-t3 · checker [email protected]

Tool Use Quality

Did the agent use the right tools, look in the right places, and avoid redundant calls?

WARNINGTool 'get_document' called 3 times in a row with identical input
run jpm-t0 · checker [email protected]

Informational

INFORecovered from 1 failed step(s) and completed
run aapl-t3
1 claim checked. Show evidence
  • 503 from document store (retried) w1
INFORecovered from 1 failed step(s) and completed
run msft-t4
1 claim checked. Show evidence
  • 503 from document store (retried) w1
INFORepeated 'get_document' failures (2× in this run); possible lost-navigation (no call to it succeeded)
run tsla-t3
2 claims checked. Show evidence
  • timeout fetching document e1
  • timeout fetching document e2
INFORetried a failing 'get_document' call unchanged: same input as the failed call at step e1
run tsla-t3
2 claims checked. Show evidence
  • timeout fetching document e1
  • identical retry e2

Evaluation provenance