The world’s most advanced language, image, and video models on one high-concurrency OpenAI-compatible API.
Licensed upstream routing — not a scraped consumer-account pool. Peak capacity, measured latency, and contract SLA are listed separately. US and Singapore are live. We do not serve PRC-registered entities.
Paid accounts default to 500 RPM; 1000+ RPM is available on request. No enterprise contract required to use the same infrastructure — start at $5 PAYG.
curl https://api.aimodelapi.ai/v1/chat/completions \
-H "Authorization: Bearer sk-ama-YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"claude-sonnet-5","messages":[{"role":"user","content":"ping"}]}'Performance data comes from ModelAPI production or standardized tests. Unless an enterprise contract says otherwise, peak throughput, latency, and cache-hit rates are historical measurements or platform ceilings — they do not mean every model, region, account, or workload will sustain them. Actual performance depends on model, input length, concurrency, network location, upstream capacity, and cache policy. See method
Trending This Week
Based on real request volume over the past 7 days · updated hourly
Harness, session budgets, and compaction live on /harness and /enterprise — not on the homepage.
Sound familiar?
Account friction, cross-border routing, and billing — we offer a compliant SaaS path (see /compliance).
Overseas phone + credit card required. Accounts banned without warning; balance lost.
✓ Our solution
Sign up with email. Eligible users per /compliance; credits never expire.
Cross-border API calls over long routes suffer latency spikes and timeouts — hard to meet production SLAs.
✓ Our solution
Multi-region PoPs and smart routing. Primary endpoint in US East; regional BASE_URLs as PoPs ship.
Foreign card only. No RMB payment, no VAT invoice, no corporate billing.
✓ Our solution
Eligible accounts (see /compliance) can pay with WeChat, Alipay, or Stripe. Invoices are for eligible non-PRC entities — we do not serve mainland-registered companies.
Ten keys, inflated bills, flaky routing — condensed into one control plane.
OpenAI, Anthropic, Google, DeepSeek — each with its own account, billing, key rotation, and quota tracking.
Per-token API billing adds up. Some gateways add markup, inflate counts, or swap models.
Trading cost vs quality manually breaks unattended workflows.
Non-developers get stuck; ticket cycles are long.
Same credential surface for experimentation and scaled traffic.
Claude Opus 5, Sonnet 5, GPT-5.6 Sol/Terra/Luna, Wan 2.7, Gemini 3.1 Pro — same integration surface.
DeepSeek V4 Flash, Qwen, Kimi, GLM — built for volume and everyday workloads.
Representative routing: code pane, streamed body, latency and credits.
Loading...Ecosystem
Point Cursor, Windsurf, VS Code, Claude Code, or Codex CLI at ModelAPI.
Universal config — works in every tool
Resilience
Reuse the exact same key — swap hostname only.
Catalog
58+ callable models · one key. Qwen3.8 Max (qwen3.8-max) and GLM-5.2 (glm-5.2) are live.
Four outcomes
Harness, MCP, and compaction stay on developer and enterprise pages.
Up to 200M tokens/min platform throughput and 1,000+ RPM capacity. Claude Opus measured TTFT P95 under 7s in Singapore. Enterprise contracts can add 99.9%/99.95% availability — those three are not the same promise.
Claude, GPT, Gemini, DeepSeek, Qwen, Kimi, GLM, plus image and video models on one OpenAI-compatible key. Licensed upstream routing — not a scraped consumer-account pool. Catalog size is the callable list on /models.
Swap base_url from OpenAI or OpenRouter. Claude Code needs ANTHROPIC_BASE_URL. Same SDK, prompts, and tools.
List prices, live gateway probes, and a public trust center. We tell you which regions are live vs planned.
WeChat Pay, Alipay, cards, Apple Pay, Google Pay, PayPal, and USDT. Official token prices plus a 5% top-up fee.
Start in under 60 seconds
Email signup auto-creates a key. $1 test credit is granted after admin review. No card.
Get free API keyCopy curl or run the request on the onboarding page. Top up only after it works.
See curlPricing
Upstream model passes through with a spelled-out infra surcharge — zero surprise fees.
For exploration and testing
Official API prices · zero markup · min $5 top-up · no monthly seat fee
Dedicated capacity and contract SLA
Representative workflows
Illustrative personas and workflows, not named customer endorsements. Verified case studies will be published with customer permission.
"Used to juggle 7–8 API accounts. Now one ModelAPI key covers everything, billing is crystal clear, and switching between Claude Opus and DeepSeek takes zero effort."
"The Singapore PoP is legitimately fast. I was skeptical about a gateway adding latency, but median response time is actually lower than hitting the OpenAI API directly from SEA. Solid routing."
"We needed one gateway for multiple model providers. ModelAPI unified billing and routing; WeChat Pay works for our eligible accounts. Deployed to production in one afternoon."
"The PPP pricing for India is real — I'm paying ~40% less than teammates in the US for the same models. Switching my existing codebase took literally 2 lines of code."
"Looked at a lot of gateways — ModelAPI is one of the few with truly 0% per-token markup and a transparent fee structure. The nexus/cheapest route cut our batch inference cost substantially."
"Building multi-tenant where users have different model preferences. The Compare page saved hours of manual benchmarking. BYOK at $0.50/M flat is the most cost-effective routing I've found."
Recent Updates
Alibaba Wan 3.0 (wan3.0-video) is on Bailian: 480P/720P/1080P, 2–30s, billed at 65% of Singapore list. Prime: WAN-video-3.0-prime.
Alibaba Qwen3.8 Max is callable on Bailian. Public ID qwen3.8-max (1M context). Qwen3.7 Max remains the previous SKU.
Call claude-opus-5 or claude-opus-latest. Latest Grok (grok-4.6) and Claude Fable 5 are listed as pending.
Use claude-opus-5 or claude-opus-latest. Save cost with claude-opus-5-rsc. Hermes dropdown "Opus 4 7" may fail — see docs.
Now routing via nexus/cheapest. $0.14/M input, 1M context.
Pause agentic workflows mid-execution and inject human decisions via the API.
Prompt-level cache hits on repeated context now reduce latency by ~40% at no extra cost.
No credit card. $1 welcome credit after admin review.