Request a dedicated capacity trial

Tell us the models, TPM, RPM, concurrency, and region you need. We reply with a capacity plan within one business day. Enterprise pricing is more than a token discount: reserved capacity, SLA, priority support, monthly minimums, and prepaid volume.

Actual quota depends on model, region, account tier, and upstream capacity. Peak figures are not the default for every account.

Developer

Individuals and test projects

Shared quota: 20 RPM free / 500 RPM after top-up

Best effort

Business

SaaS, agents, and automation

Higher quota and priority routing. 1000+ RPM on request — not automatic after signup.

Platform service objective (not contract credits)

Enterprise

Large apps and production systems

Contracted dedicated TPM/RPM

99.9% / 99.95% monthly availability by contract

Performance and capacity

Platform throughput capacity up to 200M tokens/min. 1,000+ RPM platform capacity.

Reserved TPM

Contracted tokens/min for named models — not the 200M platform ceiling.

Reserved RPM

Dedicated request rate. Paid default is 500 RPM; 1,000+ is on request.

Burst concurrency

Headroom above reserved RPM for short spikes, subject to upstream.

Dedicated model pools

Named capacity for Claude Opus, GPT-5.6 Sol, and other flagship IDs.

Cross-region routing

US and Singapore live; additional regions by plan.

Automatic failover

Multi-path licensed upstream routing — not a scraped account pool.

Priority queue

Enterprise traffic is scheduled ahead of best-effort shared quota.

Reliability

  • · 99.9% / 99.95% monthly availability SLA
  • · Public status page
  • · Incident notification
  • · Service credits in contract
  • · Named technical support
  • · Capacity early-warning

Governance

  • · Teams and sub-accounts
  • · API key spend caps
  • · Model allowlists
  • · Usage export
  • · Audit logs
  • · Auto-refill or monthly invoice
  • · Invoices and enterprise contract

Security and compliance

  • · Prompt and completion retention limits
  • · We do not train ModelAPI models on API content
  • · Upstream provider data policies apply
  • · Region routing options
  • · Encrypted API keys at rest
  • · Abuse detection and access control

Prompts are routed to upstream providers to fulfill your request. We do not use API content to train ModelAPI models. Limited logs are kept for billing disputes, abuse investigation, and legal compliance.

Capacity request

Loading form…

Performance data comes from ModelAPI production or standardized tests. Unless an enterprise contract says otherwise, peak throughput, latency, and cache-hit rates are historical measurements or platform ceilings — they do not mean every model, region, account, or workload will sustain them. Actual performance depends on model, input length, concurrency, network location, upstream capacity, and cache policy. /sla