Verdict
Quality matched on 93% of cases. Review the mismatches before changing models.
Cases evaluated
Evaluation complete
Quality matches
93%
Est. monthly savings
Projected from matching events observed over 30 days
Status
The experiment finished and its verdict is available.
Controlled replay compares both models from the same worker. Production latency is shown separately as context and is not subtracted from candidate replay latency.
gpt-5.5 and gpt-4o received the same frozen prompts in alternating order.
Replayed source p50
2,079 ms
gpt-5.5
Replayed candidate p50
946 ms
gpt-4o
Paired median delta
-1,102 ms (-54%)
95% interval: -1,222 ms to -1,033 ms
Controlled p90
2,284 ms / 1,055 ms
Source / candidate
Production source context
1,194 ms p50, 2,898 ms p90
30 captured measurements from the user's environment
Controlled coverage
30 successful pairs
Only complete source/candidate pairs enter the comparison (provider-completion-v1)
Replay failures
The evaluation compares structured content canonically and reports presentation adherence as separate evidence unless strict formatting is enabled for the feature.
This run's encrypted evidence expires at the date shown above, 14 days after the run was created. The verdict summary and savings estimate remain in experiment history after that run-specific evidence is deleted.
Review the case results before opening sensitive prompt and response evidence. Opening a case is audited.
| Case | Result | Format | Evaluation confidence | Attempts | Production source | Replayed source | Candidate | Replay delta | Status | Case details |
|---|---|---|---|---|---|---|---|---|---|---|
| 11 | Equivalent | Met | 97% | 1 | 371 ms 28,543 in / 5,763 out | 2,330 ms 28,543 in / 5,763 out | 880 ms 27,401 in / 4,841 out | -1,450 ms (-62%) | Evaluated |
0 source, 0 candidate
Failures are excluded from percentiles and retained as reliability evidence
| View details |
| 12 | Equivalent | Met | 97% | 2 | 1,473 ms 47,571 in / 8,935 out | 1,867 ms 47,571 in / 8,935 out | 911 ms 45,668 in / 7,505 out | -956 ms (-51%) | Evaluated | View details |
| 13 | Equivalent | Met | 97% | 1 | 1,028 ms 140,550 in / 23,231 out | 1,914 ms 140,550 in / 23,231 out | 942 ms 134,928 in / 19,514 out | -972 ms (-51%) | Evaluated | View details |
| 14 | Equivalent | Met | 97% | 1 | 1,268 ms 54,058 in / 8,712 out | 1,961 ms 54,058 in / 8,712 out | 973 ms 51,896 in / 7,318 out | -988 ms (-50%) | Evaluated | View details |
| 15 | Equivalent | Met | 97% | 1 | 2,788 ms 137,739 in / 20,908 out | 2,008 ms 137,739 in / 20,908 out | 1,004 ms 132,229 in / 17,563 out | -1,004 ms (-50%) | Evaluated | View details |
| 16 | Equivalent | Met | 97% | 1 | 1,523 ms 237,855 in / 40,208 out | 2,055 ms 237,855 in / 40,208 out | 1,035 ms 228,341 in / 33,775 out | -1,020 ms (-50%) | Evaluated | View details |
| 17 | Equivalent | Met | 97% | 1 | 2,378 ms 123,684 in / 25,555 out | 2,102 ms 123,684 in / 25,555 out | 1,066 ms 118,737 in / 21,466 out | -1,036 ms (-49%) | Evaluated | Hide details |
Case 17 evidenceFull captured evidence for this scored comparison. EquivalentFormat metscored97% evaluation confidenceJudge: gemini-3.1-pro-preview Evaluation notesThe candidate preserved the original answer's policy, resolution, and customer-safe guidance. Format evidenceThe response preserved the expected support-answer structure. Captured prompt | ||||||||||
| 18 | Equivalent | Met | 97% | 1 | 4,602 ms 232,449 in / 45,793 out | 2,149 ms 232,449 in / 45,793 out | 837 ms 223,151 in / 38,466 out | -1,312 ms (-61%) | Evaluated | View details |
| 19 | Equivalent | Met | 97% | 1 | 534 ms 84,763 in / 14,654 out | 2,196 ms 84,763 in / 14,654 out | 868 ms 81,372 in / 12,309 out | -1,328 ms (-60%) | Evaluated | View details |
| 20 | Equivalent | Met | 97% | 1 | 3,885 ms 264,884 in / 44,676 out | 2,243 ms 264,884 in / 44,676 out | 899 ms 254,289 in / 37,528 out | -1,344 ms (-60%) | Evaluated | View details |
Northstar support case 17: resolve the conversation-recap request with the account policy and cited help-center context.
gpt-5.5
Resolved conversation-recap case 17 with the correct policy, next step, and customer-safe explanation.
gpt-4o
Cheaper candidate preserved the policy, action, and explanation for conversation-recap case 17.