Compare lower-cost models against sampled production traffic before changing what your product runs.
Eligible traffic is sampled automatically. Each run compares one source model with a lower-cost candidate.
Sampling
The SDK must mark each event as eligible for prompt and response capture.
Latency replay
When opted in, the source model runs again beside the candidate for a controlled comparison.
Selection
trAIce selects an eligible high-volume workload and a compatible candidate that costs less.
Verdict
A full verdict evaluates 30 fresh cases. The same comparison can run again after seven days.
Measured workloads
trAIce shows up to the six highest-volume eligible feature and source-model pairs. A pair must use a supported priced model with a lower-cost candidate and have current-format, complete, unexpired prompt and response samples. Synthetic simulator workspaces compare models within the simulator provider. A run starts with 30 usable pairs.
Monthly budget
$25
Spent
$2.22
Reserved
$0
Remaining
Active and completed runs in one history. Verdict summaries remain after sensitive case evidence expires.
| Experiment | Progress | Verdict | Estimated monthly savings | Status | Completed or started | |
|---|---|---|---|---|---|---|
chat-resolver gpt-5.5 → gpt-4o-mini Judge: gemini-3.1-pro-preview (29), gemini-3.1-flash-lite fallback (1) OpenAI · Connected OpenAI key | 30 of 30 cases evaluated | Switch recommended 29 quality matches (97%) Latency p50: 1,860 ms source, 540 ms candidate (30 pairs) Format: 29 of 30 met (presentation separate) | $16,148/mo· 96.3% lower | Done |
Each captured prompt and response pair expires up to 14 days after it is created. Starting a run copies 30 selected pairs into encrypted evidence with a separate 14-day expiry. Turn sampling off at any time to delete both captured samples and retained evidence immediately.
Owners and admins can allow trAIce to retain redacted text prompt and output pairs. Each captured pair expires up to 14 days after it is created. Starting a run copies 30 selected pairs into encrypted evidence with a separate 14-day expiry. Eligible samples are sent to Google Vertex AI and the selected candidate provider. System instructions, tools, retrieval context, attachments, and agent state are not captured and are not supported by these experiments. Opt out to delete retained samples and evidence immediately.
Off by default. When enabled, each experiment sends the same 30 frozen prompts to the original source provider again so source and candidate latency can be compared from the same worker. This can double model generation calls and adds source-provider spend. The new source output is discarded after timing and usage are recorded. The captured production response remains the quality reference. Opting out stops future source legs in an active run after any in-flight provider request finishes.
By default, experiments compare parsed JSON content and report presentation adherence separately. Add an exact feature name only when whitespace, code fences, or another presentation detail is part of the real downstream contract. Changes apply to future experiment runs.
SDK must also be configured to send prompt/output on event payloads. See docs.
| Workload | Status | Fresh samples | Fresh for next run | Next expiry | Workload actions |
|---|---|---|---|---|---|
chat-resolver gpt-5.5 | Done | 6 fresh, 30 needed | 6 | Sep 7, 04:33 AM | |
conversation-recap gpt-5.5 | Done | 6 fresh, 30 needed | 6 | Sep 7, 04:33 AM | |
knowledge-agent gpt-5.5 | Ready for experiment | 36 fresh, 30 needed | 36 | Sep 7, 04:33 AM | |
ticket-summary claude-sonnet-4-6 | Done | 6 fresh, 30 needed | 6 | Sep 7, 04:33 AM |
$22.78
Evaluation costs draw from this budget. Candidate requests use a connected provider key when available.
Included experiment funding resets monthly. Next reset: Sep 1, 12:00 AM.
| Aug 23, 04:33 AM |
| View details |
ticket-summary claude-sonnet-4-6 → claude-haiku-4-5 Judge: gemini-3.1-pro-preview (29), gemini-3.1-flash-lite fallback (1) Anthropic · Connected Anthropic key | 30 of 30 cases evaluated | Switch recommended 30 quality matches (100%) Latency p50: 1,420 ms source, 610 ms candidate (30 pairs) Format: 30 of 30 met (presentation separate) | $4,092/mo· 81.8% lower | Done | Aug 22, 04:33 AM | View details |
conversation-recap gpt-5.5 → gpt-4o Judge: gemini-3.1-pro-preview (29), gemini-3.1-flash-lite fallback (1) OpenAI · Connected OpenAI key | 30 of 30 cases evaluated | Review before switching 28 quality matches (93%) Latency p50: 2,080 ms source, 930 ms candidate (30 pairs) Format: 28 of 30 met (presentation separate) | $4,466/mo· 70.8% lower | Done | Aug 20, 04:33 AM | View details |