Official token prices · zero markup · 5% infra fee only · below industry average

Official pricing, pass-through billingLower total cost than typical gateways

Transparent billing at official API prices — zero per-token markup. Only a 5% infrastructure fee at top-up (bandwidth, global PoPs, failover). Credits never expire. Read /compliance for eligibility before signing up.

0% per-token markup5% infra fee onlyWeChat · Alipay · StripeMin $5 top-upLicensed B2B channels

Infrastructure Fee — Lower as You Grow

Charged once at top-up. Covers bandwidth, global PoP infrastructure, and failover routing. Tokens billed at official API prices — 0% per-token markup.

Starter
5%
one-time infra fee
Most users
0% per-token markup
Growth
4%
one-time infra fee
$500+/month
0% per-token markup
Enterprise
3%
one-time infra fee
Committed vol.
0% per-token markup
This 5% covers global bandwidth, routing, failover, billing, and local payment rails. Tokens pass through at official prices with zero markup.

Nexus Smart Routing — Stable Aliases, Auto Model Pick

Call nexus/* and our routing algorithm picks from live models by goal (cost, coding, reasoning, speed, etc.). Each request resolves to one model at official pricing. Providers not yet connected show as Coming Soon on the Models page.

Smart Auto
Best quality/cost balance among all live models
Typical resolution
Smart algorithm pick
nexus/autoBilled at resolved model rate
Cheapest
Lowest-cost active model in the catalog
Typical resolution
e.g. deepseek-v4-flash
nexus/cheapestBilled at resolved model rate
Cheapest CN
Cheapest model from China-based providers
Typical resolution
e.g. deepseek-v4-flash
nexus/cheapest-cnBilled at resolved model rate
Cheapest West
Cheapest model from Western providers
Typical resolution
Varies by catalog
nexus/cheapest-westBilled at resolved model rate
Best Coding
Strongest coding-capable model available
Typical resolution
Highest code-tier model
nexus/best-codingBilled at resolved model rate
Best Reasoning
Best reasoning model under $5/M input
Typical resolution
Top reasoning model
nexus/best-reasoningBilled at resolved model rate
Fastest
Fastest flash/turbo/mini model available
Typical resolution
e.g. deepseek-v4-flash
nexus/fastestBilled at resolved model rate
Long Context
Cheapest model with 500K+ context window
Typical resolution
e.g. gemini-3.1-flash-lite
nexus/long-contextBilled at resolved model rate
Web Search
Perplexity Sonar with live web search
Typical resolution
e.g. sonar-pro
nexus/web-searchBilled at resolved model rate
Billing shows the resolved model ID and token usage. This is not a multi-model pipeline — each request calls one model.

Choose Your Plan

All plans are prepaid. Credits never expire.

Prices for:

Free

$0

Free models, no credit needed

  • 20 rpm · 200 rpd
  • 10+ free models
  • GLM-5 Flash · Gemini Flash
  • Basic analytics
  • Community support
Start Free
MOST POPULAR

Pay-as-you-go

$5 min top-up

Official token prices · 5% infra fee · Nexus routing aliases

Min top-up $5 · ¥35 · ₹420

  • 500 rpm · 50K rpd
  • All 50+ models
  • 0% per-token markup
  • 5% infra fee (below industry avg)
  • 8 Nexus routing aliases
  • Bring Your Own Key · $0.50/M flat
  • WeChat · Alipay · Stripe · Apple Pay
Top Up Now

Enterprise

Custom

High volume · negotiated service fee · SLA guarantee

  • Unlimited rate limits
  • Reduced service fee (negotiated)
  • Dedicated infrastructure
  • 99.9% SLA guarantee
  • 24/7 dedicated support
  • VAT invoice
  • Custom contracts
Contact Sales

ModelAPI vs Official Provider Pricing

ModelAPI bills at official API prices (0% per-token markup). The 5% service fee is charged once at credit purchase. BYOK: bring your own key, pay $0.50/M routing only — best for expensive models.

ModelOfficial output /MModelAPI output /MBYOK /MNote
Claude Fable 5$50.00$50.00= official$0.505th-gen flagship · AKL/BYOK
Claude Sonnet 5$9.50$9.50= official$0.50Intro $2/$10 · claude-sonnet-latest
Claude Opus 4.8$23.75$23.75= official$0.50Live · claude-opus-latest or BR/AKL
Claude Opus 4.7$23.75$23.75= official$0.500% markup · BYOK saves 98%
GPT-5.5$10.00$10.00= official$0.500% markup · BYOK saves 98%
GPT-5.4$10.00$10.00= official$0.500% markup · BYOK saves 97%
Gemini 3.1 Pro$12.00$12.00= official$0.500% markup · BYOK saves 96%
Grok 4.3$2.50$2.50= official0% markup · needs XAI_API_KEY
DeepSeek V4 Flash$0.28$0.28= official0% markup · platform DeepSeek key

ModelAPI vs Others

Compared to similar API gateway SaaS — multi-region endpoints, local payments for eligible users, and Nexus routing.

FeatureModelAPIOthers
Per-token markup0%0% (passthrough)
Service fee on credit buy5%5–6%
Min credit purchase$5 (¥35)$5–$10
Nexus routing aliases✓ 8 presets✗ 无
WeChat / Alipay✓ 支持✗ 不支持
Multi-region API endpoints✓ 支持✗ 不支持
Own API Key routing fee$0.50/M 固定模型价格 5–10%
PPP pricing (India/SEA)✓ 35–45% off✗ 无
Total models50+50–300+

Full Plan Comparison

FeatureFreePay-as-you-goEnterprise
Free tier
Models10 free modelsAll 50+ models50+ + custom
Per-token markup0%0%0%
Credit service fee5%2%
Min credit purchase$5 / ¥35Custom
Rate limits500 rpm / 50K rpd500 rpm / 50K rpdUnlimited
Nexus routing aliases
Bring Your Own Key · $0.50/M
Auto fallback routing
Batch API (50% off)
Usage analyticsBasicFullCustom dashboards
SLA guarantee99.9% uptime
Dedicated infrastructure
WeChat / Alipay payment
增值税专用发票
Multi-region API endpoints
SupportCommunityPriority email24/7 dedicated

FAQ

Subscription vs raw API vs ModelAPI + Harness

Claude Code Max is ~$100/mo for Opus-level usage; raw API bills per token and agents can burn $70+ overnight. Harness adds bulk routing, session budgets, and stop-loss so pay-as-you-go stays competitive — we do not claim it is always cheaper than a subscription, but it prevents runaway spend.

FactorClaude subscriptionRaw API (no Harness)ModelAPI + Harness
Monthly cost predictabilityFixed ~$100/mo (Max tier)Unbounded — users report $70+ in one nightSession budget + hourly cap → HTTP 429
Bulk task routingAll tasks on subscription modelYou pick model per request — easy to overpayAuto bulk → nexus/cheapest-west
Stop-loss on runaway agentsUsage caps within Anthropic productAccount spend cap only (optional)Per-session USD budget + alerts
Multi-vendor modelsAnthropic models onlyYes — but you manage routingPlan on reasoning · bulk on cheap · one API key
Honest fitBest if you stay within Max limitsBest with discipline + monitoringPay-as-you-go competitive when Harness is on
Shipped today Not a token relay

Image & Video Generation — Cost Awareness

User feedback: image/video work is the top token consumer. ModelAPI routes multimodal models via Chat Completions; dedicated media APIs vary by provider. Best practices below — honest about what ships today.

Image generation via chat models

Models with vision/output capabilities (e.g. Gemini, GPT-4o) can return images in agent workflows. Billed as chat completion tokens — often the highest token burn in a session.

Video & media — job-based workflow

Dedicated video generation APIs are provider-specific. We recommend async job patterns: submit → poll → download. Set a Harness session budget before long media runs.

Harness budget tips for media

Use x-harness-session-id + x-harness-budget-usd on every media batch. Enable stop-loss hourly cap in the dashboard. Route planning to reasoning models, asset prep to bulk tier.

What exists today: POST /v1/images/generations is live (Seedream, Qwen/Wan, GPT Image 2, Gemini Image). See /models?capability=image.Model catalog →