Verdict
Quality matched on 97% of cases and the candidate is cheaper. Latency is informational during calibration.
Cases evaluated
Evaluation complete
Quality matches
97%
Est. monthly savings
Projected from matching events observed over 30 days
Status
The experiment finished and its verdict is available.
Controlled replay compares both models from the same worker. Production latency is shown separately as context and is not subtracted from candidate replay latency.
gpt-5.5 and gpt-4o-mini received the same frozen prompts in alternating order.
Replayed source p50
1,859 ms
gpt-5.5
Replayed candidate p50
556 ms
gpt-4o-mini
Paired median delta
-1,272 ms (-70%)
95% interval: -1,384 ms to -1,200 ms
Controlled p90
2,064 ms / 665 ms
Source / candidate
Production source context
1,287 ms p50, 2,950 ms p90
30 captured measurements from the user's environment
Controlled coverage
30 successful pairs
Only complete source/candidate pairs enter the comparison (provider-completion-v1)
Replay failures
The evaluation compares structured content canonically and reports presentation adherence as separate evidence unless strict formatting is enabled for the feature.
This run's encrypted evidence expires at the date shown above, 14 days after the run was created. The verdict summary and savings estimate remain in experiment history after that run-specific evidence is deleted.
Review the case results before opening sensitive prompt and response evidence. Opening a case is audited.
| Case | Result | Format | Evaluation confidence | Attempts | Production source | Replayed source | Candidate | Replay delta | Status | Case details |
|---|---|---|---|---|---|---|---|---|---|---|
| 21 | Equivalent | Met | 97% | 1 | 694 ms 29,722 in / 1,153 out | 2,070 ms 29,722 in / 1,153 out | 540 ms 28,533 in / 969 out | -1,530 ms (-74%) | Evaluated |
0 source, 0 candidate
Failures are excluded from percentiles and retained as reliability evidence
| View details |
| 22 | Equivalent | Met | 97% | 1 | 1,588 ms 122,662 in / 4,298 out | 2,117 ms 122,662 in / 4,298 out | 571 ms 117,756 in / 3,610 out | -1,546 ms (-73%) | Evaluated | View details |
| 23 | Equivalent | Met | 97% | 2 | 1,257 ms 58,972 in / 1,742 out | 1,654 ms 58,972 in / 1,742 out | 602 ms 56,613 in / 1,463 out | -1,052 ms (-64%) | Evaluated | View details |
| 24 | Equivalent | Met | 97% | 1 | 595 ms 137,995 in / 4,530 out | 1,701 ms 137,995 in / 4,530 out | 633 ms 132,475 in / 3,805 out | -1,068 ms (-63%) | Evaluated | View details |
| 25 | Equivalent | Met | 97% | 1 | 2,936 ms 235,889 in / 8,712 out | 1,748 ms 235,889 in / 8,712 out | 664 ms 226,453 in / 7,318 out | -1,084 ms (-62%) | Evaluated | View details |
| 26 | Equivalent | Met | 97% | 1 | 2,070 ms 86,807 in / 2,788 out | 1,795 ms 86,807 in / 2,788 out | 695 ms 83,335 in / 2,342 out | -1,100 ms (-61%) | Evaluated | View details |
| 27 | Equivalent | Met | 97% | 1 | 281 ms 29,014 in / 1,045 out | 1,842 ms 29,014 in / 1,045 out | 466 ms 27,853 in / 878 out | -1,376 ms (-75%) | Evaluated | View details |
| 28 | Equivalent | Met | 97% | 1 | 2,502 ms 265,375 in / 9,382 out | 1,889 ms 265,375 in / 9,382 out | 497 ms 254,760 in / 7,881 out | -1,392 ms (-74%) | Evaluated | View details |
| 29 | Equivalent | Met | 97% | 1 | 636 ms 57,793 in / 1,966 out | 1,936 ms 57,793 in / 1,966 out | 528 ms 55,481 in / 1,651 out | -1,408 ms (-73%) | Evaluated | View details |
| 30 | Not equivalent | Missed | 91% | 1 | 2,355 ms 134,928 in / 5,111 out | 1,983 ms 134,928 in / 5,111 out | 559 ms 129,531 in / 4,293 out | -1,424 ms (-72%) | Evaluated | View details |