← All four sample reports

Trust Report — pm-morning-briefing-agent

Agent: atlas-briefing-agent · 40 runs evaluated · generated 2026-08-16 02:30 UTC

Gate: PASSED Dimensions with a contract threshold pass at violation_rate <= threshold; all other dimensions are zero-tolerance for violations. A dimension whose checks were skipped on every run, or for which no checker ran at all, is unevaluated and blocks a clean pass.

Findings are what Agent TrustKit's checks could verify against the contract and the evidence available in the trace, not an exhaustive audit. A PASSED gate means nothing checked failed; it does not certify that no other issue exists.

Trust profile

DimensionViolation rateThreshold ViolationsWarningsStatus
task completion 2.5% (1/40 runs) ≤ 5% 1 0 pass
groundedness 0.0% (0/40 runs) ≤ 0% 0 0 pass
completeness 0.0% (0/40 runs) ≤ 0% 0 1 pass
constraint adherence 0.0% (0/40 runs) ≤ 0% 0 0 pass
tool use quality 0.0% (0/40 runs) ≤ 5% 0 1 pass
efficiency 2.5% (1/40 runs) ≤ 5% 1 0 pass
recovery escalation 2.5% (1/40 runs) ≤ 5% 1 0 pass

Efficiency

MetricMeanMedian P95Max
Cost per run (USD)$0.0400 $0.0256 $0.0344 $0.6100
Latency per run (ms)3328 3234 4400 4735

Tool reliability

Per-tool execution reliability across the runs: how often each tool was called and how often the call errored. Deterministic from tool-call error signals, over every observed tool. Informational. Never gates.

Tool calls: 46 · failures: 4 (8.7%)

ToolInvocations FailuresFailure rate
get_document45 4 8.9%
web_lookup1 0 0.0%

Cross-trial consistency

TaskTrials Trials w/ violations Tool-set similarityStep-count CV Cost CVDuration CV
brief-nvda ⚠ mixed 51 0.80 0.12 1.61 0.20
brief-tsla ⚠ mixed 51 1.00 0.00 0.55 0.18
brief-aapl 50 1.00 0.12 0.11 0.25
brief-amzn 50 1.00 0.00 0.20 0.23
brief-goog 50 1.00 0.00 0.13 0.22
brief-jpm 50 1.00 0.24 0.19 0.04
brief-meta 50 1.00 0.00 0.10 0.19
brief-msft 50 1.00 0.12 0.20 0.17

⚠ 2 of 8 tasks are mixed: some trials violate, some don't. A pooled violation rate hides this; mixed tasks are where the agent is unpredictable.

Informational. Never gates. Tool-set similarity: 1.00 means identical tools every trial. CV = stdev/mean; higher = less repeatable.

Scope

Findings

Efficiency

Did the agent stay within cost, latency, and step budgets?

VIOLATIONRun cost $0.6100 exceeded budget $0.5000
run nvda-t2 · checker [email protected]

Task Completion

Did the agent achieve the stated objective and produce the required output?

VIOLATIONRun produced no final output
run tsla-t3 · checker [email protected]

Recovery Escalation

When steps failed, did the agent recover or escalate rather than spiral?

VIOLATIONRun ended in an unrecovered error after 2 failed step(s)
run tsla-t3 · checker [email protected]

Tool Use Quality

Did the agent use the right tools, look in the right places, and avoid redundant calls?

WARNINGTool 'get_document' called 3 times in a row with identical input
run jpm-t0 · checker [email protected]

Completeness

Did the output cover everything the sources and task required, with nothing missed?

WARNINGOutput only partially covers: "Margin mix is a watch item"
run amzn-t1 · checker [email protected]+prompts-0.12 · analyst claude-opus-4-8

Informational

INFORecovered from 1 failed step(s) and completed
run aapl-t3
1 claim checked. Show evidence
  • 503 from document store (retried) w1
INFORecovered from 1 failed step(s) and completed
run msft-t4
1 claim checked. Show evidence
  • 503 from document store (retried) w1
INFORepeated 'get_document' failures (2× in this run); possible lost-navigation (no call to it succeeded)
run tsla-t3
2 claims checked. Show evidence
  • timeout fetching document e1
  • timeout fetching document e2
INFORetried a failing 'get_document' call unchanged: same input as the failed call at step e1
run tsla-t3
2 claims checked. Show evidence
  • timeout fetching document e1
  • identical retry e2
INFOGroundedness: 2 claims checked against 4 sources; 2 supported, 0 contradicted, 0 unverifiable
run aapl-t0
2 claims checked. Show evidence
  • supported: Demand trends are in line with guidance. · “demand trends in line with guidance” s1/note-1 primary evidence
  • supported: No rating changes were published this week. · “No rating changes were published” s1/note-3 primary evidence
INFOCompleteness: 2 expected facts (derived from sources); 2 conveyed, 0 missing
run aapl-t0
2 claims checked. Show evidence
  • covered: Demand trends in line with guidance · “demand trends in line with guidance” primary evidence
  • covered: Margin mix is a watch item · “margin mix as a watch item” primary evidence
INFOGroundedness: 2 claims checked against 4 sources; 2 supported, 0 contradicted, 0 unverifiable
run msft-t2
2 claims checked. Show evidence
  • supported: Demand trends are in line with guidance. · “demand trends in line with guidance” s1/note-1 primary evidence
  • supported: No rating changes were published this week. · “No rating changes were published” s1/note-3 primary evidence
INFOCompleteness: 3 expected facts (human-authored); 3 conveyed, 0 missing
run msft-t2
3 claims checked. Show evidence
  • covered: Demand trends in line with guidance · “demand trends in line with guidance” primary evidence
  • covered: Margin mix flagged as a watch item into earnings · “margin mix as a watch item into earnings” primary evidence
  • covered: No rating changes this week · “No rating changes were published on MSFT this week” primary evidence
INFOGroundedness: 2 claims checked against 4 sources; 2 supported, 0 contradicted, 0 unverifiable
run nvda-t0
2 claims checked. Show evidence
  • supported: Demand trends are in line with guidance. · “demand trends in line with guidance” s1/note-1 primary evidence
  • supported: No rating changes were published this week. · “No rating changes were published” s1/note-3 primary evidence
INFOCompleteness: 2 expected facts (derived from sources); 2 conveyed, 0 missing
run nvda-t0
2 claims checked. Show evidence
  • covered: Demand trends in line with guidance · “demand trends in line with guidance” primary evidence
  • covered: Margin mix is a watch item · “margin mix as a watch item” primary evidence
INFOGroundedness: 2 claims checked against 4 sources; 2 supported, 0 contradicted, 0 unverifiable
run amzn-t1
2 claims checked. Show evidence
  • supported: Demand trends are in line with guidance. · “demand trends in line with guidance” s1/note-1 primary evidence
  • supported: No rating changes were published this week. · “No rating changes were published” s1/note-3 primary evidence
INFOCompleteness: 2 expected facts (derived from sources); 2 conveyed, 0 missing
run amzn-t1
2 claims checked. Show evidence
  • covered: Demand trends in line with guidance · “demand in line with guidance” primary evidence
  • partial: Margin mix is a watch item · “A desk note mentions margins” primary evidence
INFOGroundedness: 2 claims checked against 4 sources; 2 supported, 0 contradicted, 0 unverifiable
run goog-t4
2 claims checked. Show evidence
  • supported: Demand trends are in line with guidance. · “demand trends in line with guidance” s1/note-1 primary evidence
  • supported: No rating changes were published this week. · “No rating changes were published” s1/note-3 primary evidence
INFOCompleteness: 2 expected facts (derived from sources); 2 conveyed, 0 missing
run goog-t4
2 claims checked. Show evidence
  • covered: Demand trends in line with guidance · “demand trends in line with guidance” primary evidence
  • covered: Margin mix is a watch item · “margin mix as a watch item” primary evidence
INFOGroundedness: 2 claims checked against 4 sources; 2 supported, 0 contradicted, 0 unverifiable
run meta-t3
2 claims checked. Show evidence
  • supported: Demand trends are in line with guidance. · “demand trends in line with guidance” s1/note-1 primary evidence
  • supported: No rating changes were published this week. · “No rating changes were published” s1/note-3 primary evidence
INFOCompleteness: 2 expected facts (derived from sources); 2 conveyed, 0 missing
run meta-t3
2 claims checked. Show evidence
  • covered: Demand trends in line with guidance · “demand trends in line with guidance” primary evidence
  • covered: Margin mix is a watch item · “margin mix as a watch item” primary evidence
INFOGroundedness: 2 claims checked against 4 sources; 2 supported, 0 contradicted, 0 unverifiable
run jpm-t2
2 claims checked. Show evidence
  • supported: Demand trends are in line with guidance. · “demand trends in line with guidance” s1/note-1 primary evidence
  • supported: No rating changes were published this week. · “No rating changes were published” s1/note-3 primary evidence
INFOCompleteness: 2 expected facts (derived from sources); 2 conveyed, 0 missing
run jpm-t2
2 claims checked. Show evidence
  • covered: Demand trends in line with guidance · “demand trends in line with guidance” primary evidence
  • covered: Margin mix is a watch item · “margin mix as a watch item” primary evidence

Evidence kind: primary = out-of-band ground truth (retrieved sources, authoritative git diff); self-reported = the agent's own tool stream, which can corroborate a claim but never condemn it; no evidence = nothing in the trace speaks to the claim (unverifiable, which is a capture-gap signal, not a defect).

Evaluation provenance