Trust Report — pm-morning-briefing-agent
Agent: atlas-briefing-agent ·
40 runs evaluated ·
generated 2026-08-16 02:30 UTC
Gate: FAILED
— groundedness, completeness
Dimensions with a contract threshold pass at violation_rate <= threshold; all other dimensions are zero-tolerance for violations. A dimension whose checks were skipped on every run, or for which no checker ran at all, is unevaluated and blocks a clean pass.
Findings are what Agent TrustKit's checks could verify against the contract and the evidence available in the trace, not an exhaustive audit. A PASSED gate means nothing checked failed; it does not certify that no other issue exists.
Trust profile
| Dimension | Violation rate | Threshold |
Violations | Warnings | Status |
| task completion |
2.5%
(1/40 runs) |
≤ 5% |
1 |
0 |
pass |
| groundedness |
5.0%
(2/40 runs) |
≤ 0% |
2 |
0 |
fail |
| completeness |
2.5%
(1/40 runs) |
≤ 0% |
1 |
1 |
fail |
| constraint adherence |
0.0%
(0/40 runs) |
≤ 0% |
0 |
0 |
pass |
| tool use quality |
0.0%
(0/40 runs) |
≤ 5% |
0 |
1 |
pass |
| efficiency |
5.0%
(2/40 runs) |
≤ 5% |
2 |
0 |
pass |
| recovery escalation |
2.5%
(1/40 runs) |
≤ 5% |
1 |
0 |
pass |
Efficiency
| Metric | Mean | Median |
P95 | Max |
| Cost per run (USD) | $0.0387 |
$0.0251 |
$0.0348 |
$0.6100 |
| Latency per run (ms) | 4907 |
3504 |
4635 |
61200 |
Tool reliability
Per-tool execution reliability across the runs: how often
each tool was called and how often the call errored. Deterministic from
tool-call error signals, over every observed tool. Informational. Never gates.
Tool calls: 49 · failures:
4
(8.2%)
| Tool | Invocations |
Failures | Failure rate |
| get_document | 45 |
4 |
8.9% |
| web_lookup | 4 |
0 |
0.0% |
Cross-trial consistency
| Task | Trials |
Trials w/ violations |
Tool-set similarity | Step-count CV |
Cost CV | Duration CV |
| brief-aapl ⚠ mixed |
5 | 1 |
0.80 |
0.14 |
0.27 |
0.18 |
| brief-goog ⚠ mixed |
5 | 1 |
0.80 |
0.12 |
0.20 |
0.20 |
| brief-meta ⚠ mixed |
5 | 1 |
1.00 |
0.00 |
0.15 |
1.52 |
| brief-msft ⚠ mixed |
5 | 1 |
0.80 |
0.14 |
0.20 |
0.23 |
| brief-nvda ⚠ mixed |
5 | 1 |
0.80 |
0.12 |
1.63 |
0.26 |
| brief-tsla ⚠ mixed |
5 | 1 |
1.00 |
0.00 |
0.53 |
0.12 |
| brief-amzn |
5 | 0 |
1.00 |
0.00 |
0.30 |
0.15 |
| brief-jpm |
5 | 0 |
1.00 |
0.24 |
0.13 |
0.17 |
⚠ 6 of 8 tasks are mixed:
some trials violate, some don't. A pooled violation rate hides this; mixed tasks are
where the agent is unpredictable.
Informational. Never gates. Tool-set similarity: 1.00 means identical
tools every trial. CV = stdev/mean; higher = less repeatable.
Failure-associated patterns
Behaviors seen disproportionately in violating runs. Descriptive
co-occurrence over this run set. Not causes, and small counts are weak evidence.
- tool sequence 'get_document → web_lookup' · 3/6 violating runs vs
1/34 clean runs
(lift 17.0×)
- used tool 'web_lookup' · 3/6 violating runs vs
1/34 clean runs
(lift 17.0×)
Scope
- Daily internal morning briefings from the research corpus
- Coverage-universe tickers (8 names)
- Client-facing distribution (not approved)
- Trade recommendations or execution (not approved)
- Names outside the covered universe (not approved)
Findings
Efficiency
Did the agent stay within cost, latency, and step budgets?
VIOLATIONRun cost $0.6100 exceeded budget $0.5000
VIOLATIONRun took 61200ms, over the 60000ms budget
Task Completion
Did the agent achieve the stated objective and produce the required output?
VIOLATIONRun produced no final output
Recovery Escalation
When steps failed, did the agent recover or escalate rather than spiral?
VIOLATIONRun ended in an unrecovered error after 2 failed step(s)
- step
e2 · timeout fetching document
Tool Use Quality
Did the agent use the right tools, look in the right places, and avoid redundant calls?
WARNINGTool 'get_document' called 3 times in a row with identical input
- step
r0 · first of 3 identical calls, input={'id': 'jpm-note-1'}
Groundedness
Are the claims in the output supported by the sources the agent retrieved?
VIOLATIONContradicted claim in output: "Two desks downgraded AAPL to Sell overnight."
- step
s1/aapl-note-3 · “No rating changes were published on AAPL this week.” primary evidence
VIOLATIONContradicted claim in output: "Management pre-announced a $4B buyback acceleration."
- step
s1/goog-note-4 · “No capital-return or buyback announcements were made on GOOG this quarter.” primary evidence
Completeness
Did the output cover everything the sources and task required, with nothing missed?
VIOLATIONOutput omits material fact: "Margin mix flagged as a watch item into earnings"
- expected fact (human-authored) primary evidence
WARNINGOutput only partially covers: "Margin mix is a watch item"
- “A desk note mentions margins” primary evidence
Informational
INFORecovered from 1 failed step(s) and completed
run aapl-t3
1 claim checked. Show evidence
- 503 from document store (retried)
w1
INFORecovered from 1 failed step(s) and completed
run msft-t4
1 claim checked. Show evidence
- 503 from document store (retried)
w1
INFORepeated 'get_document' failures (2× in this run); possible lost-navigation (no call to it succeeded)
run tsla-t3
2 claims checked. Show evidence
- timeout fetching document
e1
- timeout fetching document
e2
INFORetried a failing 'get_document' call unchanged: same input as the failed call at step e1
run tsla-t3
2 claims checked. Show evidence
- timeout fetching document
e1
- identical retry
e2
INFOGroundedness: 2 claims checked against 4 sources; 1 supported, 1 contradicted, 0 unverifiable
run aapl-t0
2 claims checked. Show evidence
- supported: Demand trends are in line with guidance. · “demand trends in line with guidance”
s1/aapl-note-1 primary evidence
- contradicted: Two desks downgraded AAPL to Sell overnight. · “No rating changes were published on AAPL this week.”
s1/aapl-note-3 primary evidence
INFOCompleteness: 2 expected facts (derived from sources); 2 conveyed, 0 missing
run aapl-t0
2 claims checked. Show evidence
- covered: Demand trends in line with guidance · “demand trends in line with guidance” primary evidence
- covered: Margin mix is a watch item · “margin mix as a watch item” primary evidence
INFOGroundedness: 2 claims checked against 4 sources; 2 supported, 0 contradicted, 0 unverifiable
run msft-t2
2 claims checked. Show evidence
- supported: Demand trends are in line with guidance. · “demand trends in line with guidance”
s1/note-1 primary evidence
- supported: No rating changes were published this week. · “No rating changes were published”
s1/note-3 primary evidence
INFOCompleteness: 3 expected facts (human-authored); 2 conveyed, 1 missing
run msft-t2
3 claims checked. Show evidence
- covered: Demand trends in line with guidance · “demand trends in line with guidance” primary evidence
- missing: Margin mix flagged as a watch item into earnings primary evidence
- covered: No rating changes this week · “No rating changes were published on MSFT this week” primary evidence
INFOGroundedness: 2 claims checked against 4 sources; 2 supported, 0 contradicted, 0 unverifiable
run nvda-t0
2 claims checked. Show evidence
- supported: Demand trends are in line with guidance. · “demand trends in line with guidance”
s1/note-1 primary evidence
- supported: No rating changes were published this week. · “No rating changes were published”
s1/note-3 primary evidence
INFOCompleteness: 2 expected facts (derived from sources); 2 conveyed, 0 missing
run nvda-t0
2 claims checked. Show evidence
- covered: Demand trends in line with guidance · “demand trends in line with guidance” primary evidence
- covered: Margin mix is a watch item · “margin mix as a watch item” primary evidence
INFOGroundedness: 2 claims checked against 4 sources; 2 supported, 0 contradicted, 0 unverifiable
run amzn-t1
2 claims checked. Show evidence
- supported: Demand trends are in line with guidance. · “demand trends in line with guidance”
s1/note-1 primary evidence
- supported: No rating changes were published this week. · “No rating changes were published”
s1/note-3 primary evidence
INFOCompleteness: 2 expected facts (derived from sources); 2 conveyed, 0 missing
run amzn-t1
2 claims checked. Show evidence
- covered: Demand trends in line with guidance · “demand in line with guidance” primary evidence
- partial: Margin mix is a watch item · “A desk note mentions margins” primary evidence
INFOGroundedness: 2 claims checked against 4 sources; 1 supported, 1 contradicted, 0 unverifiable
run goog-t4
2 claims checked. Show evidence
- supported: Demand trends are in line with guidance. · “demand trends in line with guidance”
s1/goog-note-1 primary evidence
- contradicted: Management pre-announced a $4B buyback acceleration. · “No capital-return or buyback announcements were made on GOOG this quarter.”
s1/goog-note-4 primary evidence
INFOCompleteness: 2 expected facts (derived from sources); 2 conveyed, 0 missing
run goog-t4
2 claims checked. Show evidence
- covered: Demand trends in line with guidance · “demand trends in line with guidance” primary evidence
- covered: Margin mix is a watch item · “margin mix as a watch item” primary evidence
INFOGroundedness: 2 claims checked against 4 sources; 2 supported, 0 contradicted, 0 unverifiable
run meta-t3
2 claims checked. Show evidence
- supported: Demand trends are in line with guidance. · “demand trends in line with guidance”
s1/note-1 primary evidence
- supported: No rating changes were published this week. · “No rating changes were published”
s1/note-3 primary evidence
INFOCompleteness: 2 expected facts (derived from sources); 2 conveyed, 0 missing
run meta-t3
2 claims checked. Show evidence
- covered: Demand trends in line with guidance · “demand trends in line with guidance” primary evidence
- covered: Margin mix is a watch item · “margin mix as a watch item” primary evidence
INFOGroundedness: 2 claims checked against 4 sources; 2 supported, 0 contradicted, 0 unverifiable
run jpm-t2
2 claims checked. Show evidence
- supported: Demand trends are in line with guidance. · “demand trends in line with guidance”
s1/note-1 primary evidence
- supported: No rating changes were published this week. · “No rating changes were published”
s1/note-3 primary evidence
INFOCompleteness: 2 expected facts (derived from sources); 2 conveyed, 0 missing
run jpm-t2
2 claims checked. Show evidence
- covered: Demand trends in line with guidance · “demand trends in line with guidance” primary evidence
- covered: Margin mix is a watch item · “margin mix as a watch item” primary evidence
Evidence kind: primary = out-of-band ground truth (retrieved sources, authoritative git diff); self-reported = the agent's own tool stream, which can corroborate a claim but never condemn it; no evidence = nothing in the trace speaks to the claim (unverifiable, which is a capture-gap signal, not a defect).
Evaluation provenance