Verdict
Quality matched on 93% of cases. Review the mismatches before changing models.
Cases evaluated
Evaluation complete
Quality matches
93%
Est. monthly savings
Projected from matching events observed over 30 days
Status
The experiment finished and its verdict is available.
Controlled replay compares both models from the same worker. Production latency is shown separately as context and is not subtracted from candidate replay latency.
gpt-5.5 and gpt-4o received the same frozen prompts in alternating order.
Replayed source p50
2,079 ms
gpt-5.5
Replayed candidate p50
946 ms
gpt-4o
Paired median delta
-1,102 ms (-54%)
95% interval: -1,222 ms to -1,033 ms
Controlled p90
2,284 ms / 1,055 ms
Source / candidate
Production source context
1,194 ms p50, 2,898 ms p90
30 captured measurements from the user's environment
Controlled coverage
30 successful pairs
Only complete source/candidate pairs enter the comparison (provider-completion-v1)
Replay failures
The evaluation compares structured content canonically and reports presentation adherence as separate evidence unless strict formatting is enabled for the feature.
This run's encrypted evidence expires at the date shown above, 14 days after the run was created. The verdict summary and savings estimate remain in experiment history after that run-specific evidence is deleted.
Review the case results before opening sensitive prompt and response evidence. Opening a case is audited.
| Case | Result | Format | Evaluation confidence | Attempts | Production source | Replayed source | Candidate | Replay delta | Status | Case details |
|---|---|---|---|---|---|---|---|---|---|---|
| 21 | Equivalent | Met | 97% | 1 | 1,854 ms 83,033 in / 13,224 out | 2,290 ms 83,033 in / 13,224 out | 930 ms 79,712 in / 11,108 out | -1,360 ms (-59%) | Evaluated |
0 source, 0 candidate
Failures are excluded from percentiles and retained as reliability evidence
| View details |
| 22 | Equivalent | Met | 97% | 1 | 185 ms 27,894 in / 4,959 out | 2,337 ms 27,894 in / 4,959 out | 961 ms 26,778 in / 4,166 out | -1,376 ms (-59%) | Evaluated | View details |
| 23 | Equivalent | Met | 97% | 2 | 1,546 ms 74,384 in / 12,867 out | 1,874 ms 74,384 in / 12,867 out | 992 ms 71,409 in / 10,808 out | -882 ms (-47%) | Evaluated | View details |
| 24 | Equivalent | Met | 97% | 1 | 771 ms 27,245 in / 5,629 out | 1,921 ms 27,245 in / 5,629 out | 1,023 ms 26,155 in / 4,728 out | -898 ms (-47%) | Evaluated | View details |
| 25 | Equivalent | Met | 97% | 1 | 635 ms 31,137 in / 5,495 out | 1,968 ms 31,137 in / 5,495 out | 1,054 ms 29,892 in / 4,616 out | -914 ms (-46%) | Evaluated | View details |
| 26 | Equivalent | Met | 97% | 1 | 1,415 ms 49,733 in / 8,935 out | 2,015 ms 49,733 in / 8,935 out | 1,085 ms 47,744 in / 7,505 out | -930 ms (-46%) | Evaluated | View details |
| 27 | Equivalent | Met | 97% | 1 | 795 ms 48,652 in / 8,042 out | 2,062 ms 48,652 in / 8,042 out | 856 ms 46,706 in / 6,755 out | -1,206 ms (-58%) | Evaluated | Hide details |
Case 27 evidenceFull captured evidence for this scored comparison. EquivalentFormat metscored97% evaluation confidenceJudge: gemini-3.1-pro-preview Evaluation notesThe candidate preserved the original answer's policy, resolution, and customer-safe guidance. Format evidenceThe response preserved the expected support-answer structure. Captured prompt | ||||||||||
| 28 | Equivalent | Met | 97% | 1 | 2,672 ms 112,440 in / 20,908 out | 2,109 ms 112,440 in / 20,908 out | 887 ms 107,942 in / 17,563 out | -1,222 ms (-58%) | Evaluated | View details |
| 29 | Not equivalent | Missed | 91% | 1 | 1,432 ms 140,550 in / 23,812 out | 2,156 ms 140,550 in / 23,812 out | 918 ms 134,928 in / 20,002 out | -1,238 ms (-57%) | Evaluated | View details |
| 30 | Not equivalent | Missed | 91% | 1 | 4,400 ms 243,260 in / 45,793 out | 2,203 ms 243,260 in / 45,793 out | 949 ms 233,530 in / 38,466 out | -1,254 ms (-57%) | Evaluated | View details |
Northstar support case 27: resolve the conversation-recap request with the account policy and cited help-center context.
gpt-5.5
Resolved conversation-recap case 27 with the correct policy, next step, and customer-safe explanation.
gpt-4o
Cheaper candidate preserved the policy, action, and explanation for conversation-recap case 27.