Request a dedicated capacity trial
Tell us the models, TPM, RPM, concurrency, and region you need. We reply with a capacity plan within one business day. Enterprise pricing is more than a token discount: reserved capacity, SLA, priority support, monthly minimums, and prepaid volume.
Actual quota depends on model, region, account tier, and upstream capacity. Peak figures are not the default for every account.
Developer
Individuals and test projects
Shared quota: 20 RPM free / 500 RPM after top-up
Best effort
Business
SaaS, agents, and automation
Higher quota and priority routing. 1000+ RPM on request — not automatic after signup.
Platform service objective (not contract credits)
Enterprise
Large apps and production systems
Contracted dedicated TPM/RPM
99.9% / 99.95% monthly availability by contract
Performance and capacity
Platform throughput capacity up to 200M tokens/min. 1,000+ RPM platform capacity.
Reserved TPM
Contracted tokens/min for named models — not the 200M platform ceiling.
Reserved RPM
Dedicated request rate. Paid default is 500 RPM; 1,000+ is on request.
Burst concurrency
Headroom above reserved RPM for short spikes, subject to upstream.
Dedicated model pools
Named capacity for Claude Opus, GPT-5.6 Sol, and other flagship IDs.
Cross-region routing
US and Singapore live; additional regions by plan.
Automatic failover
Multi-path licensed upstream routing — not a scraped account pool.
Priority queue
Enterprise traffic is scheduled ahead of best-effort shared quota.
Reliability
- · 99.9% / 99.95% monthly availability SLA
- · Public status page
- · Incident notification
- · Service credits in contract
- · Named technical support
- · Capacity early-warning
Governance
- · Teams and sub-accounts
- · API key spend caps
- · Model allowlists
- · Usage export
- · Audit logs
- · Auto-refill or monthly invoice
- · Invoices and enterprise contract
Security and compliance
- · Prompt and completion retention limits
- · We do not train ModelAPI models on API content
- · Upstream provider data policies apply
- · Region routing options
- · Encrypted API keys at rest
- · Abuse detection and access control
Prompts are routed to upstream providers to fulfill your request. We do not use API content to train ModelAPI models. Limited logs are kept for billing disputes, abuse investigation, and legal compliance.
Capacity request
Loading form…
Performance data comes from ModelAPI production or standardized tests. Unless an enterprise contract says otherwise, peak throughput, latency, and cache-hit rates are historical measurements or platform ceilings — they do not mean every model, region, account, or workload will sustain them. Actual performance depends on model, input length, concurrency, network location, upstream capacity, and cache policy. /sla