Verdict
Quality matched on 97% of cases and the candidate is cheaper. Latency is informational during calibration.
Cases evaluated
Evaluation complete
Quality matches
97%
Est. monthly savings
Projected from matching events observed over 30 days
Status
The experiment finished and its verdict is available.
Controlled replay compares both models from the same worker. Production latency is shown separately as context and is not subtracted from candidate replay latency.
gpt-5.5 and gpt-4o-mini received the same frozen prompts in alternating order.
Replayed source p50
1,859 ms
gpt-5.5
Replayed candidate p50
556 ms
gpt-4o-mini
Paired median delta
-1,272 ms (-70%)
95% interval: -1,384 ms to -1,200 ms
Controlled p90
2,064 ms / 665 ms
Source / candidate
Production source context
1,287 ms p50, 2,950 ms p90
30 captured measurements from the user's environment
Controlled coverage
30 successful pairs
Only complete source/candidate pairs enter the comparison (provider-completion-v1)
Replay failures
The evaluation compares structured content canonically and reports presentation adherence as separate evidence unless strict formatting is enabled for the feature.
This run's encrypted evidence expires at the date shown above, 14 days after the run was created. The verdict summary and savings estimate remain in experiment history after that run-specific evidence is deleted.
Review the case results before opening sensitive prompt and response evidence. Opening a case is audited.
| Case | Result | Format | Evaluation confidence | Attempts | Production source | Replayed source | Candidate | Replay delta | Status | Case details |
|---|---|---|---|---|---|---|---|---|---|---|
| 11 | Equivalent | Met | 97% | 1 | 108 ms 28,307 in / 965 out | 2,110 ms 28,307 in / 965 out | 490 ms 27,175 in / 811 out | -1,620 ms (-77%) | Evaluated |
0 source, 0 candidate
Failures are excluded from percentiles and retained as reliability evidence
| View details |
| 12 | Equivalent | Met | 97% | 2 | 997 ms 48,357 in / 1,698 out | 1,647 ms 48,357 in / 1,698 out | 521 ms 46,423 in / 1,426 out | -1,126 ms (-68%) | Evaluated | View details |
| 13 | Equivalent | Met | 97% | 1 | 3,076 ms 144,128 in / 4,414 out | 1,694 ms 144,128 in / 4,414 out | 552 ms 138,363 in / 3,708 out | -1,142 ms (-67%) | Evaluated | View details |
| 14 | Equivalent | Met | 97% | 1 | 2,027 ms 247,683 in / 8,488 out | 1,741 ms 247,683 in / 8,488 out | 583 ms 237,776 in / 7,130 out | -1,158 ms (-67%) | Evaluated | View details |
| 15 | Equivalent | Met | 97% | 1 | 1,681 ms 90,581 in / 2,716 out | 1,788 ms 90,581 in / 2,716 out | 614 ms 86,958 in / 2,281 out | -1,174 ms (-66%) | Evaluated | View details |
| 16 | Equivalent | Met | 97% | 1 | 873 ms 54,254 in / 1,832 out | 1,835 ms 54,254 in / 1,832 out | 645 ms 52,084 in / 1,539 out | -1,190 ms (-65%) | Evaluated | View details |
| 17 | Equivalent | Met | 97% | 1 | 377 ms 47,178 in / 1,921 out | 1,882 ms 47,178 in / 1,921 out | 676 ms 45,291 in / 1,614 out | -1,206 ms (-64%) | Evaluated | View details |
| 18 | Equivalent | Met | 97% | 1 | 1,836 ms 141,062 in / 4,995 out | 1,929 ms 141,062 in / 4,995 out | 447 ms 135,420 in / 4,196 out | -1,482 ms (-77%) | Evaluated | View details |
| 19 | Equivalent | Met | 97% | 1 | 5,107 ms 241,786 in / 9,605 out | 1,976 ms 241,786 in / 9,605 out | 478 ms 232,115 in / 8,068 out | -1,498 ms (-76%) | Evaluated | Hide details |
Case 19 evidenceFull captured evidence for this scored comparison. EquivalentFormat metscored97% evaluation confidenceJudge: gemini-3.1-pro-preview Evaluation notesThe candidate preserved the original answer's policy, resolution, and customer-safe guidance. Format evidenceThe response preserved the expected support-answer structure. Captured prompt | ||||||||||
| 20 | Equivalent | Met | 97% | 1 | 750 ms 88,694 in / 3,074 out | 2,023 ms 88,694 in / 3,074 out | 509 ms 85,146 in / 2,582 out | -1,514 ms (-75%) | Evaluated | View details |
Northstar support case 19: resolve the chat-resolver request with the account policy and cited help-center context.
gpt-5.5
Resolved chat-resolver case 19 with the correct policy, next step, and customer-safe explanation.
gpt-4o-mini
Cheaper candidate preserved the policy, action, and explanation for chat-resolver case 19.