GPT
GPT-5.6 measured latency
GPT-5.6 Sol, Terra, and Luna are in the public catalog. Model-level TTFT P50/P95 history is not yet published as audit data. What is public today: gateway probes plus the list-price board.
| Model | TTFT P50 | TTFT P95 | Status |
|---|---|---|---|
| gpt-5.6-sol | Awaiting probes | Awaiting probes | Callable in catalog |
| gpt-5.6-terra | Awaiting probes | Awaiting probes | Callable in catalog |
| gpt-5.6-luna | Awaiting probes | Awaiting probes | Callable in catalog |
Self-serve check: /tools/api-latency-test. List prices: /benchmarks. Gateway probe latency is not GPT time-to-first-token.
Performance data comes from ModelAPI production or standardized tests. Unless an enterprise contract says otherwise, peak throughput, latency, and cache-hit rates are historical measurements or platform ceilings — they do not mean every model, region, account, or workload will sustain them. Actual performance depends on model, input length, concurrency, network location, upstream capacity, and cache policy.