Verdict
Quality matched on 97% of cases and the candidate is cheaper. Latency is informational during calibration.
Cases evaluated
Evaluation complete
Quality matches
97%
Est. monthly savings
Projected from matching events observed over 30 days
Status
The experiment finished and its verdict is available.
Controlled replay compares both models from the same worker. Production latency is shown separately as context and is not subtracted from candidate replay latency.
gpt-5.5 and gpt-4o-mini received the same frozen prompts in alternating order.
Replayed source p50
1,859 ms
gpt-5.5
Replayed candidate p50
556 ms
gpt-4o-mini
Paired median delta
-1,272 ms (-70%)
95% interval: -1,384 ms to -1,200 ms
Controlled p90
2,064 ms / 665 ms
Source / candidate
Production source context
1,287 ms p50, 2,950 ms p90
30 captured measurements from the user's environment
Controlled coverage
30 successful pairs
Only complete source/candidate pairs enter the comparison (provider-completion-v1)
Replay failures
The evaluation compares structured content canonically and reports presentation adherence as separate evidence unless strict formatting is enabled for the feature.
This run's encrypted evidence expires at the date shown above, 14 days after the run was created. The verdict summary and savings estimate remain in experiment history after that run-specific evidence is deleted.
Review the case results before opening sensitive prompt and response evidence. Opening a case is audited.
| Case | Result | Format | Evaluation confidence | Attempts | Production source | Replayed source | Candidate | Replay delta | Status | Case details |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Equivalent | Met | 97% | 2 | 478 ms 49,537 in / 1,921 out | 1,640 ms 49,537 in / 1,921 out | 440 ms 47,556 in / 1,614 out | -1,200 ms (-73%) | Evaluated |
0 source, 0 candidate
Failures are excluded from percentiles and retained as reliability evidence
| View details |
| 2 | Equivalent | Met | 97% | 1 | 797 ms 144,128 in / 4,530 out | 1,687 ms 144,128 in / 4,530 out | 471 ms 138,363 in / 3,805 out | -1,216 ms (-72%) | Evaluated | View details |
| 3 | Equivalent | Met | 97% | 1 | 1,357 ms 51,896 in / 1,832 out | 1,734 ms 51,896 in / 1,832 out | 502 ms 49,820 in / 1,539 out | -1,232 ms (-71%) | Evaluated | View details |
| 4 | Equivalent | Met | 97% | 1 | 1,119 ms 241,786 in / 9,829 out | 1,781 ms 241,786 in / 9,829 out | 533 ms 232,115 in / 8,256 out | -1,248 ms (-70%) | Evaluated | View details |
| 5 | Equivalent | Met | 97% | 1 | 737 ms 50,716 in / 1,653 out | 1,828 ms 50,716 in / 1,653 out | 564 ms 48,687 in / 1,389 out | -1,264 ms (-69%) | Evaluated | View details |
| 6 | Equivalent | Met | 97% | 1 | 2,557 ms 150,261 in / 4,298 out | 1,875 ms 150,261 in / 4,298 out | 595 ms 144,251 in / 3,610 out | -1,280 ms (-68%) | Evaluated | Hide details |
Case 6 evidenceFull captured evidence for this scored comparison. EquivalentFormat metscored97% evaluation confidenceJudge: gemini-3.1-pro-preview Evaluation notesThe candidate preserved the original answer's policy, resolution, and customer-safe guidance. Format evidenceThe response preserved the expected support-answer structure. Captured prompt | ||||||||||
| 7 | Equivalent | Met | 97% | 1 | 361 ms 86,807 in / 2,859 out | 1,922 ms 86,807 in / 2,859 out | 626 ms 83,335 in / 2,402 out | -1,296 ms (-67%) | Evaluated | View details |
| 8 | Equivalent | Met | 97% | 1 | 1,617 ms 49,537 in / 1,876 out | 1,969 ms 49,537 in / 1,876 out | 657 ms 47,556 in / 1,576 out | -1,312 ms (-67%) | Evaluated | View details |
| 9 | Equivalent | Met | 97% | 1 | 1,317 ms 147,195 in / 4,879 out | 2,016 ms 147,195 in / 4,879 out | 688 ms 141,307 in / 4,098 out | -1,328 ms (-66%) | Evaluated | View details |
| 10 | Equivalent | Met | 97% | 1 | 4,198 ms 253,581 in / 9,382 out | 2,063 ms 253,581 in / 9,382 out | 459 ms 243,438 in / 7,881 out | -1,604 ms (-78%) | Evaluated | View details |
Northstar support case 6: resolve the chat-resolver request with the account policy and cited help-center context.
gpt-5.5
Resolved chat-resolver case 6 with the correct policy, next step, and customer-safe explanation.
gpt-4o-mini
Cheaper candidate preserved the policy, action, and explanation for chat-resolver case 6.