Claude Opus

Claude Opus measured TTFT

Claude Opus TTFT P95 under 7s. Dated measured snapshot — not a contract SLA and not every account’s default.

Model
Claude Opus (catalog IDs: claude-opus-5, claude-opus-latest → Opus 5)
Metric
TTFT P95 — time to first byte, not first meaningful token
Percentile
P95
Result
under 7s
Region
Singapore
Sample
10,000
Window
2026-08-01 to 2026-08-07
Network
client network time
Retries
excluded
Fallback
excluded

Not published: input-length mix, thinking on/off, or whether the same P95 holds at peak. Model-level P50/P95 history needs a probe job writing continuously; until then this page is a method-labeled snapshot, not a live board.

Performance data comes from ModelAPI production or standardized tests. Unless an enterprise contract says otherwise, peak throughput, latency, and cache-hit rates are historical measurements or platform ceilings — they do not mean every model, region, account, or workload will sustain them. Actual performance depends on model, input length, concurrency, network location, upstream capacity, and cache policy.

/performance · claude-opus-5