1. Measure three clocks
Record queue wait, model generation and result delivery separately. Their causes and mitigations are different.
CORE FORMULAEffective cost = attempt cost ÷ acceptance rate
2. Test peak and quiet periods
A median from quiet hours does not describe production reliability. Capture p50 and p95 latency across realistic traffic windows.
3. Include failed jobs
Timeouts and retries affect both latency and cost. Report successful-only speed separately from end-to-end accepted-output time.
Models worth comparing next
Frequently asked questions
What is p95 latency?
Ninety-five percent of measured requests finish at or below that time.
How many tests are useful?
Run enough across multiple periods to expose queue variation, not just a short burst.
Next step
Build your own workload plan, then save candidate models for comparison. Your plan stays in your browser.
Open workload planner →