GPT

GPT-5.6 measured latency

GPT-5.6 Sol, Terra, and Luna are in the public catalog. Model-level TTFT P50/P95 history is not yet published as audit data. What is public today: gateway probes plus the list-price board.

ModelTTFT P50TTFT P95Status
gpt-5.6-solAwaiting probesAwaiting probesCallable in catalog
gpt-5.6-terraAwaiting probesAwaiting probesCallable in catalog
gpt-5.6-lunaAwaiting probesAwaiting probesCallable in catalog

Self-serve check: /tools/api-latency-test. List prices: /benchmarks. Gateway probe latency is not GPT time-to-first-token.

Performance data comes from ModelAPI production or standardized tests. Unless an enterprise contract says otherwise, peak throughput, latency, and cache-hit rates are historical measurements or platform ceilings — they do not mean every model, region, account, or workload will sustain them. Actual performance depends on model, input length, concurrency, network location, upstream capacity, and cache policy.