Baseline run (40 runs, gate FAILED) → candidate run_approved (40 runs, gate PASSED)
| Dimension | Baseline | Candidate | Δ | Direction |
|---|---|---|---|---|
| task completion | 2.5% | 2.5% | +0.0% | — unchanged |
| groundedness | 5.0% | 0.0% | -5.0% | ▼ improved |
| completeness | 2.5% | 0.0% | -2.5% | ▼ improved |
| constraint adherence | 0.0% | 0.0% | +0.0% | — unchanged |
| tool use quality | 0.0% | 0.0% | +0.0% | — unchanged |
| efficiency | 5.0% | 2.5% | -2.5% | ▼ improved |
| recovery escalation | 2.5% | 2.5% | +0.0% | — unchanged |
Informational — never affects the regression verdict; small samples make these too noisy to gate on.
| Signal | Baseline | Candidate | Δ |
|---|---|---|---|
| mixed tasks (some trials violate, some don't) | 6 | 2 | -4 |
| mean tool-set similarity across trials | 0.9 | 0.975 | 0.075 |
Finding diffs match on (dimension, checker, summary) and are indicative; dimension rates are exact.