← All four sample reports

Comparison — pm-morning-briefing-agent

Baseline run (40 runs, gate FAILED) → candidate run_approved (40 runs, gate PASSED)

Verdict: NO REGRESSION
DimensionBaselineCandidateΔDirection
task completion2.5%2.5%+0.0%— unchanged
groundedness5.0%0.0%-5.0%▼ improved
completeness2.5%0.0%-2.5%▼ improved
constraint adherence0.0%0.0%+0.0%— unchanged
tool use quality0.0%0.0%+0.0%— unchanged
efficiency5.0%2.5%-2.5%▼ improved
recovery escalation2.5%2.5%+0.0%— unchanged

Cross-run signals

Informational — never affects the regression verdict; small samples make these too noisy to gate on.

SignalBaselineCandidateΔ
mixed tasks (some trials violate, some don't)62-4
mean tool-set similarity across trials0.90.9750.075

Resolved since baseline

Finding diffs match on (dimension, checker, summary) and are indicative; dimension rates are exact.