Performance

Peak capacity, measured results, contract SLA

200M TPM platform capacity, 1,000+ RPM, high-reliability Claude and GPT API. Individuals pay as they go; enterprises get dedicated capacity, automatic failover, and a contract SLA.

Peak platform capacity

Platform throughput capacity up to 200M tokens/min

1,000+ RPM platform capacity

Actual quota depends on model, region, account tier, and upstream capacity. Peak figures are not the default for every account.

Measured performance

Claude Opus TTFT P95 under 7s

Region Singapore; sample 10,000; 2026-08-01 to 2026-08-07; client network time.

Not a contract guarantee. P50/P99, thinking, input length, and peak hours are separate.

Contract SLA

99.9% / 99.95%

Availability, dedicated TPM/RPM, support response, incident notice, and service credits are written into an enterprise contract. Peak throughput and measured TTFT are not SLA guarantees.

/sla
MetricCurrentWindowMethod
API success rateAwaiting public probe seriesLast 30 daysGateway probes on /status
Claude Opus TTFT P50Awaiting model-level probesLast 24 hoursPublic page currently shows gateway probes only
Claude Opus TTFT P95Measured under 7s2026-08-01 to 2026-08-07Singapore · n=10,000 · client network time · retries/fallback excluded · ops-reported
GPT TTFT P95Awaiting model-level probesLast 24 hoursSee /benchmarks/gpt
Peak platform RPM1000+Platform ceilingPeak capacity, not default quota
Cache hit rateup to 98%Repeating-prefix workloadsOn workloads with a stable system prompt and repeated prefixes, prompt-cache hit rate can reach 98%. Results vary by model, request shape, TTL, and cache_control. This is not a natural hit rate for all traffic.
Automatic fallback countAwaiting probe seriesLast 24 hoursProduction series not yet published

For repeating Agent workloads, Prompt Cache can cut input-token cost and response latency.

Performance data comes from ModelAPI production or standardized tests. Unless an enterprise contract says otherwise, peak throughput, latency, and cache-hit rates are historical measurements or platform ceilings — they do not mean every model, region, account, or workload will sustain them. Actual performance depends on model, input length, concurrency, network location, upstream capacity, and cache policy.

/benchmarks/claude-opus · /benchmarks/gpt · /status · /status/history · /enterprise