Performance
Peak capacity, measured results, contract SLA
200M TPM platform capacity, 1,000+ RPM, high-reliability Claude and GPT API. Individuals pay as they go; enterprises get dedicated capacity, automatic failover, and a contract SLA.
Peak platform capacity
Platform throughput capacity up to 200M tokens/min
1,000+ RPM platform capacity
Actual quota depends on model, region, account tier, and upstream capacity. Peak figures are not the default for every account.
Measured performance
Claude Opus TTFT P95 under 7s
Region Singapore; sample 10,000; 2026-08-01 to 2026-08-07; client network time.
Not a contract guarantee. P50/P99, thinking, input length, and peak hours are separate.
Contract SLA
99.9% / 99.95%
Availability, dedicated TPM/RPM, support response, incident notice, and service credits are written into an enterprise contract. Peak throughput and measured TTFT are not SLA guarantees.
/sla| Metric | Current | Window | Method |
|---|---|---|---|
| API success rate | Awaiting public probe series | Last 30 days | Gateway probes on /status |
| Claude Opus TTFT P50 | Awaiting model-level probes | Last 24 hours | Public page currently shows gateway probes only |
| Claude Opus TTFT P95 | Measured under 7s | 2026-08-01 to 2026-08-07 | Singapore · n=10,000 · client network time · retries/fallback excluded · ops-reported |
| GPT TTFT P95 | Awaiting model-level probes | Last 24 hours | See /benchmarks/gpt |
| Peak platform RPM | 1000+ | Platform ceiling | Peak capacity, not default quota |
| Cache hit rate | up to 98% | Repeating-prefix workloads | On workloads with a stable system prompt and repeated prefixes, prompt-cache hit rate can reach 98%. Results vary by model, request shape, TTL, and cache_control. This is not a natural hit rate for all traffic. |
| Automatic fallback count | Awaiting probe series | Last 24 hours | Production series not yet published |
For repeating Agent workloads, Prompt Cache can cut input-token cost and response latency.
Performance data comes from ModelAPI production or standardized tests. Unless an enterprise contract says otherwise, peak throughput, latency, and cache-hit rates are historical measurements or platform ceilings — they do not mean every model, region, account, or workload will sustain them. Actual performance depends on model, input length, concurrency, network location, upstream capacity, and cache policy.
/benchmarks/claude-opus · /benchmarks/gpt · /status · /status/history · /enterprise