Official pricing, pass-through billingLower total cost than typical gateways
Transparent billing at official API prices — zero per-token markup. Only a 5% infrastructure fee at top-up (bandwidth, global PoPs, failover). Credits never expire. Read /compliance for eligibility before signing up.
Infrastructure Fee — Lower as You Grow
Charged once at top-up. Covers bandwidth, global PoP infrastructure, and failover routing. Tokens billed at official API prices — 0% per-token markup.
Nexus Smart Routing — Stable Aliases, Auto Model Pick
Call nexus/* and our routing algorithm picks from live models by goal (cost, coding, reasoning, speed, etc.). Each request resolves to one model at official pricing. Providers not yet connected show as Coming Soon on the Models page.
nexus/autoBilled at resolved model ratenexus/cheapestBilled at resolved model ratenexus/cheapest-cnBilled at resolved model ratenexus/cheapest-westBilled at resolved model ratenexus/best-codingBilled at resolved model ratenexus/best-reasoningBilled at resolved model ratenexus/fastestBilled at resolved model ratenexus/long-contextBilled at resolved model ratenexus/web-searchBilled at resolved model rateChoose Your Plan
All plans are prepaid. Credits never expire.
Free
Free models, no credit needed
- 20 rpm · 200 rpd
- 10+ free models
- GLM-5 Flash · Gemini Flash
- Basic analytics
- Community support
Pay-as-you-go
Official token prices · 5% infra fee · Nexus routing aliases
Min top-up $5 · ¥35 · ₹420
- 500 rpm · 50K rpd
- All 50+ models
- 0% per-token markup
- 5% infra fee (below industry avg)
- 8 Nexus routing aliases
- Bring Your Own Key · $0.50/M flat
- WeChat · Alipay · Stripe · Apple Pay
Enterprise
High volume · negotiated service fee · SLA guarantee
- Unlimited rate limits
- Reduced service fee (negotiated)
- Dedicated infrastructure
- 99.9% SLA guarantee
- 24/7 dedicated support
- VAT invoice
- Custom contracts
ModelAPI vs Official Provider Pricing
ModelAPI bills at official API prices (0% per-token markup). The 5% service fee is charged once at credit purchase. BYOK: bring your own key, pay $0.50/M routing only — best for expensive models.
| Model | Official output /M | ModelAPI output /M | BYOK /M | Note |
|---|---|---|---|---|
Claude Fable 5 | $50.00 | $50.00= official | $0.50 | 5th-gen flagship · AKL/BYOK |
Claude Sonnet 5 | $9.50 | $9.50= official | $0.50 | Intro $2/$10 · claude-sonnet-latest |
Claude Opus 4.8 | $23.75 | $23.75= official | $0.50 | Live · claude-opus-latest or BR/AKL |
Claude Opus 4.7 | $23.75 | $23.75= official | $0.50 | 0% markup · BYOK saves 98% |
GPT-5.5 | $10.00 | $10.00= official | $0.50 | 0% markup · BYOK saves 98% |
GPT-5.4 | $10.00 | $10.00= official | $0.50 | 0% markup · BYOK saves 97% |
Gemini 3.1 Pro | $12.00 | $12.00= official | $0.50 | 0% markup · BYOK saves 96% |
Grok 4.3 | $2.50 | $2.50= official | — | 0% markup · needs XAI_API_KEY |
DeepSeek V4 Flash | $0.28 | $0.28= official | — | 0% markup · platform DeepSeek key |
ModelAPI vs Others
Compared to similar API gateway SaaS — multi-region endpoints, local payments for eligible users, and Nexus routing.
| Feature | ModelAPI | Others |
|---|---|---|
| Per-token markup | 0% | 0% (passthrough) |
| Service fee on credit buy | 5% | 5–6% |
| Min credit purchase | $5 (¥35) | $5–$10 |
| Nexus routing aliases | ✓ 8 presets | ✗ 无 |
| WeChat / Alipay | ✓ 支持 | ✗ 不支持 |
| Multi-region API endpoints | ✓ 支持 | ✗ 不支持 |
| Own API Key routing fee | $0.50/M 固定 | 模型价格 5–10% |
| PPP pricing (India/SEA) | ✓ 35–45% off | ✗ 无 |
| Total models | 50+ | 50–300+ |
Full Plan Comparison
| Feature | Free | Pay-as-you-go | Enterprise |
|---|---|---|---|
| Free tier | |||
| Models | 10 free models | All 50+ models | 50+ + custom |
| Per-token markup | 0% | 0% | 0% |
| Credit service fee | — | 5% | 2% |
| Min credit purchase | — | $5 / ¥35 | Custom |
| Rate limits | 500 rpm / 50K rpd | 500 rpm / 50K rpd | Unlimited |
| Nexus routing aliases | |||
| Bring Your Own Key · $0.50/M | |||
| Auto fallback routing | |||
| Batch API (50% off) | |||
| Usage analytics | Basic | Full | Custom dashboards |
| SLA guarantee | 99.9% uptime | ||
| Dedicated infrastructure | |||
| WeChat / Alipay payment | |||
| 增值税专用发票 | |||
| Multi-region API endpoints | |||
| Support | Community | Priority email | 24/7 dedicated |
FAQ
Subscription vs raw API vs ModelAPI + Harness
Claude Code Max is ~$100/mo for Opus-level usage; raw API bills per token and agents can burn $70+ overnight. Harness adds bulk routing, session budgets, and stop-loss so pay-as-you-go stays competitive — we do not claim it is always cheaper than a subscription, but it prevents runaway spend.
| Factor | Claude subscription | Raw API (no Harness) | ModelAPI + Harness |
|---|---|---|---|
| Monthly cost predictability | Fixed ~$100/mo (Max tier) | Unbounded — users report $70+ in one night | Session budget + hourly cap → HTTP 429 |
| Bulk task routing | All tasks on subscription model | You pick model per request — easy to overpay | Auto bulk → nexus/cheapest-west |
| Stop-loss on runaway agents | Usage caps within Anthropic product | Account spend cap only (optional) | Per-session USD budget + alerts |
| Multi-vendor models | Anthropic models only | Yes — but you manage routing | Plan on reasoning · bulk on cheap · one API key |
| Honest fit | Best if you stay within Max limits | Best with discipline + monitoring | Pay-as-you-go competitive when Harness is on |
Image & Video Generation — Cost Awareness
User feedback: image/video work is the top token consumer. ModelAPI routes multimodal models via Chat Completions; dedicated media APIs vary by provider. Best practices below — honest about what ships today.
Image generation via chat models
Models with vision/output capabilities (e.g. Gemini, GPT-4o) can return images in agent workflows. Billed as chat completion tokens — often the highest token burn in a session.
Video & media — job-based workflow
Dedicated video generation APIs are provider-specific. We recommend async job patterns: submit → poll → download. Set a Harness session budget before long media runs.
Harness budget tips for media
Use x-harness-session-id + x-harness-budget-usd on every media batch. Enable stop-loss hourly cap in the dashboard. Route planning to reasoning models, asset prep to bulk tier.