No credit card · $1 test credit after review · 58+ callable models

Enterprise-grade AI throughput

The world’s most advanced language, image, and video models on one high-concurrency OpenAI-compatible API.

Licensed upstream routing — not a scraped consumer-account pool. Peak capacity, measured latency, and contract SLA are listed separately. US and Singapore are live. We do not serve PRC-registered entities.

Paid accounts default to 500 RPM; 1000+ RPM is available on request. No enterprise contract required to use the same infrastructure — start at $5 PAYG.

curl · OpenAI compatible
curl https://api.aimodelapi.ai/v1/chat/completions \
  -H "Authorization: Bearer sk-ama-YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"claude-sonnet-5","messages":[{"role":"user","content":"ping"}]}'
200M tokens/min
Peak platform throughput (not default quota)
1,000+
Platform RPM capacity — paid default 500
P95 < 7s
Claude Opus measured TTFT (Singapore)
98%
Cache hit on repeating prefixes (not all traffic)

Performance data comes from ModelAPI production or standardized tests. Unless an enterprise contract says otherwise, peak throughput, latency, and cache-hit rates are historical measurements or platform ceilings — they do not mean every model, region, account, or workload will sustain them. Actual performance depends on model, input length, concurrency, network location, upstream capacity, and cache policy. See method

Harness, session budgets, and compaction live on /harness and /enterprise — not on the homepage.

Sound familiar?

3 blockers that stop most developers

Account friction, cross-border routing, and billing — we offer a compliant SaaS path (see /compliance).

🚫

Account rejected or banned

Overseas phone + credit card required. Accounts banned without warning; balance lost.

✓ Our solution

Sign up with email. Eligible users per /compliance; credits never expire.

📶

Unstable network & timeouts

Cross-border API calls over long routes suffer latency spikes and timeouts — hard to meet production SLAs.

✓ Our solution

Multi-region PoPs and smart routing. Primary endpoint in US East; regional BASE_URLs as PoPs ship.

💸

Payment walls & no invoices

Foreign card only. No RMB payment, no VAT invoice, no corporate billing.

✓ Our solution

Eligible accounts (see /compliance) can pay with WeChat, Alipay, or Stripe. Invoices are for eligible non-PRC entities — we do not serve mainland-registered companies.

Problems we solve

Ten keys, inflated bills, flaky routing — condensed into one control plane.

Sound familiar?

10 provider accounts · 10 API keys · 10 dashboards

OpenAI, Anthropic, Google, DeepSeek — each with its own account, billing, key rotation, and quota tracking.

ModelAPI
One key. 58+ callable models. Unified billing.
Sound familiar?

API feels 10× more expensive · hidden token inflation

Per-token API billing adds up. Some gateways add markup, inflate counts, or swap models.

ModelAPI
0% per-token markup. Official prices. What you call is what you get.
Sound familiar?

Cheap models underperform · flagship costs · manual switching hurts flow

Trading cost vs quality manually breaks unattended workflows.

ModelAPI
nexus/* routes by task; standard IDs always hit the named model.
Sound familiar?

Complex docs · hard to debug · slow support

Non-developers get stuck; ticket cycles are long.

ModelAPI
Docs tested with non-experts; AI support available 7×24.

Frontier paths and savings in one envelope

Same credential surface for experimentation and scaled traffic.

01

Frontier models

Claude Opus 5, Sonnet 5, GPT-5.6 Sol/Terra/Luna, Wan 2.7, Gemini 3.1 Pro — same integration surface.

Claude Opus 5Claude Sonnet 5GPT-5.6 SolGPT-5.6 TerraWan 2.7Gemini 3.1 Pro
02

High-value models

DeepSeek V4 Flash, Qwen, Kimi, GLM — built for volume and everyday workloads.

DeepSeek V4 FlashQwen3.6 PlusKimi K2.6GLM-5 FlashGemini 3.1 Flash-Lite

Watch a request end-to-end

Representative routing: code pane, streamed body, latency and credits.

request.ts
Anthropic
Loading...
response.json
Awaiting response…

Ecosystem

Works where you already build

Point Cursor, Windsurf, VS Code, Claude Code, or Codex CLI at ModelAPI.

All guides

Universal config — works in every tool

Base URL
https://api.aimodelapi.ai/v1
API Key
sk-ama-YOUR_KEY

Resilience

Backup hosts when routes change

Reuse the exact same key — swap hostname only.

GLGlobal (Primary gateway)PRIMARYLive
api.aimodelapi.ai/v1
Primary · nearest PoP
USVirginiaPlanned
us.api.aimodelapi.ai/v1
Dedicated hostname planned · global host already routes to US East
SGSingaporeLive
sg.api.aimodelapi.ai/v1
Asia Pacific PoP
JPTokyoPlanned
jp.api.aimodelapi.ai/v1
Japan PoP
THBangkokPlanned
th.api.aimodelapi.ai/v1
Southeast Asia
EUFrankfurtPlanned
eu.api.aimodelapi.ai/v1
Europe PoP
BKAlt Primary DomainPlanned
api2.aimodelapi.ai/v1
Backup if .ai is blocked
Switching endpoints takes one config change
Replace api.aimodelapi.ai with any backup host in your BASE_URL. Same API key. Same code.
Live Status

Four outcomes

What you get in the first five minutes

Harness, MCP, and compaction stay on developer and enterprise pages.

Peak capacity, measured latency, contract SLA

Up to 200M tokens/min platform throughput and 1,000+ RPM capacity. Claude Opus measured TTFT P95 under 7s in Singapore. Enterprise contracts can add 99.9%/99.95% availability — those three are not the same promise.

Language, image, and video models worldwide

Claude, GPT, Gemini, DeepSeek, Qwen, Kimi, GLM, plus image and video models on one OpenAI-compatible key. Licensed upstream routing — not a scraped consumer-account pool. Catalog size is the callable list on /models.

5-minute migration

Swap base_url from OpenAI or OpenRouter. Claude Code needs ANTHROPIC_BASE_URL. Same SDK, prompts, and tools.

Verifiable price and status

List prices, live gateway probes, and a public trust center. We tell you which regions are live vs planned.

Global payments

WeChat Pay, Alipay, cards, Apple Pay, Google Pay, PayPal, and USDT. Official token prices plus a 5% top-up fee.

Start in under 60 seconds

Three steps to your first call

01
🔐

Sign up, get a key

Email signup auto-creates a key. $1 test credit is granted after admin review. No card.

Get free API key
02
⚡

Run the first call

Copy curl or run the request on the onboarding page. Top up only after it works.

See curl
03
💳

Top up $5 when ready

WeChat / Alipay / Stripe. A $5 top-up starts 24-hour onboarding.

View pricing
Change 2 lines — everything else staysPython
# Before (OpenAI direct)
base_url = "https://api.openai.com/v1"
api_key = "sk-openai-..."
# After (ModelAPI — same SDK, any model)
base_url = "https://api.aimodelapi.ai/v1"
api_key = "sk-ama-YOUR_KEY"
client = OpenAI(base_url=base_url, api_key=api_key)
resp = client.chat.completions.create(
model="claude-opus-latest", # routes to Opus 5
messages=[{"role": "user", "content": "Hello"}]
)

Pricing

Usage-based tiers without guesswork

Upstream model passes through with a spelled-out infra surcharge — zero surprise fees.

Free

$0/month

For exploration and testing

  • 20 req/min · 50K req/day
  • 10+ free models
  • Community support
  • Basic analytics
Start Free
MOST POPULAR

Pay-as-you-go

5% infra fee

Official API prices · zero markup · min $5 top-up · no monthly seat fee

  • Default 500 RPM after top-up · 50,000 req/day
  • 1000+ RPM on request — not automatic
  • All 58+ listed models
  • 0% per-token markup · 5% infra fee at top-up
  • Credits never expire
  • WeChat · Alipay · cards · PayPal · USDT
Get API Key

Enterprise

Custom

Dedicated capacity and contract SLA

  • Reserved TPM / RPM by contract
  • 99.9% or 99.95% monthly availability SLA
  • Priority routing, failover, incident notice
  • Teams, key budgets, invoices, audit logs
  • Service credits defined in the contract
Request dedicated capacity

Accepted payment methods worldwide

WeChat Pay
CNY
Alipay
CNY
Visa
Card
Mastercard
Card
Apple Pay
Wallet
Google Pay
Wallet
PayPal
Global
Amex
Card
256-bit TLS encryptionPCI DSS compliant via StripeCredits never expireMulti-currency support

Representative workflows

How different teams use a unified model API

Illustrative personas and workflows, not named customer endorsements. Verified case studies will be published with customer permission.

"Used to juggle 7–8 API accounts. Now one ModelAPI key covers everything, billing is crystal clear, and switching between Claude Opus and DeepSeek takes zero effort."

CY
Chen Yuxiang
AI indie developer

"The Singapore PoP is legitimately fast. I was skeptical about a gateway adding latency, but median response time is actually lower than hitting the OpenAI API directly from SEA. Solid routing."

JK
Jason K.
Full-stack engineer, Singapore

"We needed one gateway for multiple model providers. ModelAPI unified billing and routing; WeChat Pay works for our eligible accounts. Deployed to production in one afternoon."

LX
Lin Xiaowen
Startup CTO

"The PPP pricing for India is real — I'm paying ~40% less than teammates in the US for the same models. Switching my existing codebase took literally 2 lines of code."

PM
Priya M.
ML Engineer, Bangalore

"Looked at a lot of gateways — ModelAPI is one of the few with truly 0% per-token markup and a transparent fee structure. The nexus/cheapest route cut our batch inference cost substantially."

WB
Wang Boyuan
Algorithm engineer

"Building multi-tenant where users have different model preferences. The Compare page saved hours of manual benchmarking. BYOK at $0.50/M flat is the most cost-effective routing I've found."

AT
Alex T.
SaaS founder, London

Recent Updates

What we've shipped lately

Full changelog
Live

Wan 3.0 is live — call WAN-video-3.0

Alibaba Wan 3.0 (wan3.0-video) is on Bailian: 480P/720P/1080P, 2–30s, billed at 65% of Singapore list. Prime: WAN-video-3.0-prime.

Live

Qwen3.8 Max is live — call qwen3.8-max

Alibaba Qwen3.8 Max is callable on Bailian. Public ID qwen3.8-max (1M context). Qwen3.7 Max remains the previous SKU.

Live

Claude Opus 5 live — Grok 4.6 and Fable 5 pending

Call claude-opus-5 or claude-opus-latest. Latest Grok (grok-4.6) and Claude Fable 5 are listed as pending.

Live

Claude Opus 4.8 shipped (now previous-gen — use Opus 5)

Use claude-opus-5 or claude-opus-latest. Save cost with claude-opus-5-rsc. Hermes dropdown "Opus 4 7" may fail — see docs.

New Model

DeepSeek V4 Flash added — 344% weekly surge

Now routing via nexus/cheapest. $0.14/M input, 1M context.

Feature

Human-in-the-loop tool calls now supported

Pause agentic workflows mid-execution and inject human decisions via the API.

Performance

nexus/cheapest-west routing now 40% faster with edge caching

Prompt-level cache hits on repeated context now reduce latency by ~40% at no extra cost.

Call the world’s most advanced language, image, and video models now.

No credit card. $1 welcome credit after admin review.