Verdict
Quality matched on 93% of cases. Review the mismatches before changing models.
Cases evaluated
Evaluation complete
Quality matches
93%
Est. monthly savings
Projected from matching events observed over 30 days
Status
The experiment finished and its verdict is available.
Controlled replay compares both models from the same worker. Production latency is shown separately as context and is not subtracted from candidate replay latency.
gpt-5.5 and gpt-4o received the same frozen prompts in alternating order.
Replayed source p50
2,079 ms
gpt-5.5
Replayed candidate p50
946 ms
gpt-4o
Paired median delta
-1,102 ms (-54%)
95% interval: -1,222 ms to -1,033 ms
Controlled p90
2,284 ms / 1,055 ms
Source / candidate
Production source context
1,194 ms p50, 2,898 ms p90
30 captured measurements from the user's environment
Controlled coverage
30 successful pairs
Only complete source/candidate pairs enter the comparison (provider-completion-v1)
Replay failures
The evaluation compares structured content canonically and reports presentation adherence as separate evidence unless strict formatting is enabled for the feature.
This run's encrypted evidence expires at the date shown above, 14 days after the run was created. The verdict summary and savings estimate remain in experiment history after that run-specific evidence is deleted.
Review the case results before opening sensitive prompt and response evidence. Opening a case is audited.
| Case | Result | Format | Evaluation confidence | Attempts | Production source | Replayed source | Candidate | Replay delta | Status | Case details |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Equivalent | Met | 97% | 2 | 478 ms 45,409 in / 9,605 out | 1,860 ms 45,409 in / 9,605 out | 830 ms 43,593 in / 8,068 out | -1,030 ms (-55%) | Evaluated |
0 source, 0 candidate
Failures are excluded from percentiles and retained as reliability evidence
| View details |
| 2 | Equivalent | Met | 97% | 1 | 797 ms 132,117 in / 22,651 out | 1,907 ms 132,117 in / 22,651 out | 861 ms 126,832 in / 19,027 out | -1,046 ms (-55%) | Evaluated | View details |
| 3 | Equivalent | Met | 97% | 1 | 1,119 ms 221,637 in / 49,144 out | 1,954 ms 221,637 in / 49,144 out | 892 ms 212,772 in / 41,281 out | -1,062 ms (-54%) | Evaluated | View details |
| 4 | Equivalent | Met | 97% | 1 | 361 ms 79,573 in / 14,296 out | 2,001 ms 79,573 in / 14,296 out | 923 ms 76,390 in / 12,009 out | -1,078 ms (-54%) | Evaluated | View details |
| 5 | Equivalent | Met | 97% | 1 | 108 ms 25,948 in / 4,825 out | 2,048 ms 25,948 in / 4,825 out | 954 ms 24,910 in / 4,053 out | -1,094 ms (-53%) | Evaluated | View details |
| 6 | Equivalent | Met | 97% | 1 | 873 ms 49,733 in / 9,159 out | 2,095 ms 49,733 in / 9,159 out | 985 ms 47,744 in / 7,694 out | -1,110 ms (-53%) | Evaluated | View details |
| 7 | Equivalent | Met | 97% | 1 | 1,588 ms 112,440 in / 21,489 out | 2,142 ms 112,440 in / 21,489 out | 1,016 ms 107,942 in / 18,051 out | -1,126 ms (-53%) | Evaluated | View details |
| 8 | Equivalent | Met | 97% | 1 | 2,502 ms 243,260 in / 46,910 out | 2,189 ms 243,260 in / 46,910 out | 1,047 ms 233,530 in / 39,404 out | -1,142 ms (-52%) | Evaluated | View details |
| 9 | Equivalent | Met | 97% | 1 | 954 ms 86,493 in / 13,581 out | 2,236 ms 86,493 in / 13,581 out | 1,078 ms 83,033 in / 11,408 out | -1,158 ms (-52%) | Evaluated | View details |
| 10 | Equivalent | Met | 97% | 1 | 593 ms 48,652 in / 9,829 out | 2,283 ms 48,652 in / 9,829 out | 849 ms 46,706 in / 8,256 out | -1,434 ms (-63%) | Evaluated | View details |