AI API Providers
Compare pricing, models, and features across all major AI API providers
| # | Provider | Popular Models | Input | Output | China Available | Free Tier | |
|---|---|---|---|---|---|---|---|
| #1 | OpenAI | GPT-5, GPT-5-mini | GPT-5: $1.25/M, GPT-5-mini: $0.25/M, GPT-4o: $2.50/M, GPT-4o-mini: $0.15/M | GPT-5: $10/M, GPT-5-mini: $2/M, GPT-4o: $10/M, GPT-4o-mini: $0.60/M | ❌ Proxy required (export control restrictions for China users) | New accounts receive $5 in credits valid for 3 months | View Details → |
| #2 | Anthropic Claude | Claude Sonnet 5, Claude Opus 5 | Sonnet 5 (intro): $2/M, Opus 5: $5/M, Sonnet 5/4.5: $3/M, Opus 4.8: $5/M, Haiku 4.5: $0.80/M | Sonnet 5 (intro): $10/M, Opus 5: $25/M, Sonnet 5/4.5: $15/M, Opus 4.8: $25/M, Haiku 4.5: $4/M | ❌ Proxy required (US export controls) | Limited Claude.ai free tier (web only); no API credits; Sonnet 5 limited-time pricing at $2/M until Aug 31, 2026 | View Details → |
| #2 | Azure OpenAI | GPT-4o, GPT-4o-mini | GPT-4o: $2.50/M, GPT-4o-mini: $0.15/M, o1: $15/M, o3: $10/M | GPT-4o: $10/M, GPT-4o-mini: $0.60/M, o1: $60/M, o3: $40/M | ⚠️ Available via 21Vianet (restricted model set) | View Details → | |
| #3 | OpenRouter | GPT-4o, Claude Opus/Sonnet/Haiku | $0.065/M to $15/M 取决于模型,平均 $0.50/M | $0.26/M to $75/M 取决于模型,平均 $2/M | ⚠️ Partial (not blocked in China but unstable) | View Details → | |
| #4 | DeepSeek | DeepSeek-V3, DeepSeek-R1 | DeepSeek-V3: ¥0.14/M token (≈$0.02), R1: ¥0.14/M | DeepSeek-V3: ¥0.28/M token (≈$0.04), R1: ¥0.28/M | ✅ Direct access in China | View Details → | |
| #4 | xAI Grok | Grok 4.6, Grok 4.6 Fast | Grok 4.6: $2/M tokens (fast variant $4/M) | Grok 4.6: $6/M tokens (fast variant $12/M) | ❌ Proxy required | View Details → | |
| #4 | Gemini Multimodal | Gemini Omni Flash (multimodal, 1M context, audio+video+text), Nano Banana 2 Lite (image generation, $0.034/1K tokens) | Omni Flash: $0.15/M input tokens (multimodal); Nano Banana 2 Lite: $0.034/1K image tokens; Veo 3.1: $0.10/sec video | Omni Flash: $0.60/M output tokens (multimodal); Nano Banana Pro: $0.12/1K image tokens | ❌ Proxy required (Google AI Studio not directly accessible from mainland China; requires stable proxy) | Free tier: Gemini Omni Flash 15 RPM, Veo 3.1 2 generations/day, Imagen 4 10 images/day — no credit card required | View Details → |
| #5 | Google Gemini | Gemini 2.5 Pro, Gemini 2.5 Flash | 2.5 Pro: $1.25-10/M, 2.5 Flash: $0.15-1.25/M, 2.0 Flash: $0.10/M | 2.5 Pro: $5-40/M, 2.5 Flash: $0.30-5/M, 2.0 Flash: $0.40/M | ❌ Proxy required (Google Cloud not accessible in CN) | View Details → | |
| #5 | Qwen (Alibaba) | Qwen3.8-Flash-Next, Qwen3.8-Flash | Qwen3.8-Max: ¥8/M, Qwen3.5-Max: ¥4/M, Qwen3.5-Plus: ¥2/M | Qwen3.8-Max: ¥24/M, Qwen3.5-Max: ¥12/M, Qwen3.5-Plus: ¥6/M | ✅ Direct China access (Aliyun Bailian); global via OpenRouter and HuggingFace | 1M tokens free credit for new users (90-day validity); Qwen3.5 open weights for self-hosting | View Details → |
| #6 | Perplexity AI | Sonar Pro, Sonar | Sonar Pro: $3/M, Sonar: $1/M | Sonar Pro: $15/M, Sonar: $1/M | ❌ Proxy required | View Details → | |
| #6 | Meituan LongCat | LongCat-2.0 (1.6T MoE), LongCat-2.0 INT4 | OpenRouter: $0.60/M token (48B active MoE) | OpenRouter: $2.40/M token (48B active MoE) | ✅ OpenRouter / Hugging Face self-host | View Details → | |
| #6 | Meta Model API | Muse Spark 1.1 (muse-spark-1.1) | $1.25/1M tokens | $4.25/1M tokens | ❌ Proxy required (Meta Model API public preview for US developers only; China access requires proxy) | Free tier: 60 RPM, 2M TPM; free credits on signup for US developers (public preview) | View Details → |
| #6 | Stepfun | Step-2 (1.2T multimodal flagship, text+image+audio+video, 128K context), Step-2-mini (200B multimodal, faster Step-2 variant) | Step-2: ¥6/M, Step-2-mini: ¥1/M, Step-1: ¥4/M, Step-R: ¥8/M | Step-2: ¥18/M, Step-2-mini: ¥3/M, Step-1: ¥12/M, Step-R: ¥24/M | ✅ Direct from China (platform.stepfun.com endpoint, domestic BGP routes) | New accounts receive 1,000,000 free tokens (30-day validity from registration) | View Details → |
| #7 | Exa | Search — neural semantic search across the web; 6 search types (auto / instant / fast / deep-lite / deep / deep-reasoning); returns 10 results by default, $1/1k per extra result above 10; AI page summaries $1/1k pages, Contents — full page text / highlights / summaries for known URLs; $1/1k pages per content type (text, highlights, summary billed separately) | Pure pay-as-you-go by endpoint and search-type. Search $7/1k requests (base 10 results) + $1/1k per extra result + $1/1k AI summaries; Contents $1/1k pages; Answer $5/1k; Monitors $15/1k; Deep Search $12-15/1k; Agent fixed effort $0.012-$1.00/request or usage-based $0.10/ACU + tool calls (default $5 auto cap, $20 max cap). New accounts get $20 in free credits (~2,800 searches); Free Tier adds $10/month. No subscription, no minimum spend. | Same as input. Enterprise custom volume + Zero Data Retention + SLA + postpaid invoice. | ❌ Proxy required for mainland China. Exa runs the primary API at https://api.exa.ai from a single-region US deployment. There is no documented ICP-registered mainland China endpoint as of 2026-08-30. Mainland China production traffic typically needs a proxy or relay; transpacific first-byte latency is usually 100-300 ms. Enterprise plans can negotiate custom routing and dedicated infrastructure for regulated workloads (HIPAA / GDPR data residency). | Permanent $20 signup credit (~2,800 Search calls) + automatic $10 Free Tier monthly credit refill. No credit card required, no minimum spend. Free credits apply to every endpoint. | View Details → |
| #7 | Alibaba Cloud Bailian | Qwen3.5-Max, Qwen3.5-Plus | Qwen3.5-Max: ¥4/M, Qwen3.5-Plus: ¥2/M, Qwen3.5-72B: ¥4/M, QwQ: ¥2/M | Qwen3.5-Max: ¥12/M, Qwen3.5-Plus: ¥6/M, Qwen3.5-72B: ¥12/M, QwQ: ¥8/M | ✅ Direct access in China | View Details → | |
| #8 | Baidu ERNIE (文心一言) | ERNIE 4.5 Turbo, ERNIE 4.0 Turbo | ERNIE 4.5 Turbo: $0.003/M tokens, ERNIE 4.0 Turbo: $0.012/M | ERNIE 4.5 Turbo: $0.003/M, ERNIE 4.0 Turbo: $0.012/M | ✅ Direct access in China | View Details → | |
| #8 | Vercel AI Gateway | OpenAI GPT-4o / GPT-4o-mini / GPT-5.x, Anthropic Claude Opus 4.8 / Sonnet 4 | 免费层: $5/月 AI Gateway Credits;付费层: 按量付费,Provider 官价零加价 | Token 零加价(含 BYOK);附加功能另计 | ⚠️ Proxy required (China needs VPN) | $5/month free AI Gateway Credits per team (Free-Tier eligible models only) | View Details → |
| #9 | Moonshot AI Kimi | Kimi K3, Kimi K2.7 Code | K3: ¥2/M cached, ¥20/M uncached; K2.7/K2.6: ¥1.10-¥1.30/M cached, ¥6.50/M uncached | K3: ¥100/M; K2.7/K2.6: ¥27/M | ✅ Direct access intended for Mainland China | The ¥15 new-user coupon cannot be used for Kimi K3; recharge required | View Details → |
| #9 | Mistral AI | Mistral Large 2, Mistral Small | Large 2: $2/M, Codestral: $1/M, Ministral 8B: $0.10/M | Large 2: $6/M, Codestral: $3/M, Ministral 8B: $0.10/M | ❌ Proxy required | Free API credits on signup on La Plateforme (within daily rate limits) | View Details → |
| #9 | Dify | GPT-4o / GPT-4o-mini / o1 / o3 (via OpenAI), Claude 3.5 Sonnet / Haiku / Opus (via Anthropic) | Sandbox 免费:200 GPT-4 调用 + 5MB vector 存储 + 10 文档 + 5,000 API 调用/月;Professional $59/workplace/月:100MB向量 + 100文档 + 50 工作流 | Team $159/workplace/月:500MB向量 + 500文档 + 200 工作流 + 500K 触发事件;Enterprise 合同制(SSO/私有部署/SLA) | ✅ Parent company LangGenius is China-based; native Chinese support, Alibaba Cloud one-click deploy, SOC 2/GDPR/ISO 27001 certified, full China availability | Sandbox free tier: 200 GPT-4 calls + 5MB vector storage + 10 docs + 2 workflows + 5,000 API calls/month (30-day logs) | View Details → |
| #10 | Zhipu AI GLM | GLM-5.3, GLM-5.3-Flash | GLM-5.3 (Workers AI): $1.40/M + $0.26/M cached; GLM-5.2: ¥8/M; GLM-5.3-Flash (Workers AI): $0.15/M; GLM-5: ¥4-6/M; GLM-4.7: ¥2-4/M; GLM-4.5-Air: ¥0.8/M; GLM-4.7-Flash: ¥0/M (free) | GLM-5.3 (Workers AI): $4.40/M; GLM-5.2: ¥28/M; GLM-5.3-Flash (Workers AI): $0.50/M; GLM-5: ¥18-22/M; GLM-4.7: ¥8-16/M; GLM-4.5-Air: ¥2-8/M; GLM-4.7-Flash: ¥0/M (free) | ✅ Direct access in China | View Details → | |
| #10 | Cohere | Command R+, Command R7 | Command R7: $2.50/M, Command R: $0.50/M, Embed: $0.10/M | Command R7: $10/M, Command R: $1.50/M | ❌ Proxy required | View Details → | |
| #10 | Cloudflare AI Gateway | 通过 Gateway 路由到 20+ Provider 任意模型 | 免费层: 100万次请求/月;超出 $0.20/100万次 + Provider 自身费用 | 按请求计费,不按 token 计;Gateway 本身加价 $0.20/100万请求 | ✅ Global CDN (China edge via partners) | View Details → | |
| #11 | Tencent Hunyuan | Hunyuan Hy4 preview, Hunyuan Turbo | Turbo S: ¥0.8/M tokens, Turbo: ¥1.2/M, Lite: ¥0.3/M | Turbo S: ¥0.8/M tokens, Turbo: ¥4.8/M, Lite: ¥0.3/M | ✅ Direct access in China | View Details → | |
| #11 | Fireworks AI | Llama 3.3 70B, Firefunction-v2 | Firefunction-v2: $0.90/M, Llama 3.3 70B: $0.90/M, DeepSeek-R1: $2.00/M | Firefunction-v2: $0.90/M, Llama 3.3 70B: $0.90/M, DeepSeek-R1: $2.00/M | ❌ Proxy required | View Details → | |
| #11 | Portkey | OpenAI GPT-4o / GPT-5 / GPT-5.6 family, Anthropic Claude Opus 5 / Sonnet 5 / Haiku | Free 层: 10,000 请求/月;Hobby $49/月(10万请求);Growth $249/月(100万请求);Enterprise 合同制 | 请求量计费,不收 token 路由费;BYOK 免费;日志/可观测性/Caching 按订阅档位开放 | ❌ Proxy required (portkey.ai is unstable from mainland China; self-hosted open-source edition recommended for CN teams) | 10,000 requests/month; 100K log retention; community Slack support; BYOK free; no credit card required | View Details → |
| #11 | Black Forest Labs | FLUX.2 [pro], FLUX.1.1 [pro] | FLUX.2 [pro]: $0.05/MP, FLUX.1.1 [pro]: $0.04/MP, FLUX.1 [schnell]: $0.003/MP, FLUX.1 Fill: $0.05/MP | Flat per-megapixel billing, no tier-based markup | ❌ Proxy required (AWS US-East + EU-Frankfurt deployment) | View Details → | |
| #12 | ByteDance Doubao | Doubao-Seed-2.0 (旗舰推理), Doubao-Seed-2.0-lite (高性价比) | Seed-2.0: ¥0.8/M, Seed-2.0-lite: ¥0.5/M, Seed-2.0-mini: ¥0.15/M | Seed-2.0: ¥2/M, Seed-2.0-lite: ¥1/M, Seed-2.0-mini: ¥0.6/M | ✅ Direct access in China | View Details → | |
| #12 | LiteLLM | 通过 Proxy 转发 100+ Provider 任意模型 | 开源免费自部署;LiteLLM Cloud 按量计费 | 开源免费自部署;LiteLLM Cloud 按量计费 | ✅ Open-source self-hosted, no restrictions | View Details → | |
| #12 | SiliconFlow | Qwen/Qwen3.5-Plus, Qwen/Qwen2.5-72B-Instruct | ¥0.4-2/1M tokens (Qwen2.5-7B ~¥0.4,Llama 3.3 70B ~¥2) | ¥0.4-2/1M tokens | ✅ Direct access in China | ¥1-200 free credits on signup (varies by promo), enough for 1-20M tokens | View Details → |
| #12 | Jina AI | jina-embeddings-v3, jina-embeddings-v2-base-en | Embeddings v3: $0.02/M tokens (10K free); Reranker v3: $0.018/M tokens (10K free); Reader: $0.02/M tokens (free up to 1M tokens); CLIP-v1: $0.02/M tokens | Flat per-million-token billing, no tier-based markup; batch discount of 50% available on all endpoints | ❌ Proxy required (AWS US-East + EU-Frankfurt deployment) | View Details → | |
| #12 | Arize Phoenix | OpenTelemetry 原生 trace collector,支持任意 LLM 框架, Phoenix Cloud 托管 + 自托管 (Apache-2 / Elastic-2.0) | Cloud 免费层: 10 GiB 存储/工作空间;自托管: $0(仅基础设施费) | Phoenix Cloud 付费: 按量付费(存储计费);Arize AX 企业: 合同制 | ✅ Open-source self-host + Phoenix Cloud (China SaaS access requires proxy) | Phoenix Cloud free tier: 10 GiB storage per workspace (no time limit) | View Details → |
| #13 | Block AI (b.ai) | GPT-4o, GPT-4o-mini | 主流模型 + 0-10% 加价(如 GPT-4o $2.5→$2.75/M) | 主流模型 + 0-10% 加价 | ✅ Direct access in China | View Details → | |
| #13 | Groq | Llama 3.3 70B, Llama 3.1 405B | Llama 3.3 70B: $0.59/M, DeepSeek-R1: $0.89/M, Mixtral 8x7B: $0.15/M | Llama 3.3 70B: $0.79/M, DeepSeek-R1: $0.89/M, Mixtral 8x7B: $0.15/M | ❌ Proxy required | View Details → | |
| #13 | fal.ai | FLUX.2 [pro] / [dev] / [schnell], Kling Video 2.1 / 2.0 | FLUX.2 [pro]: $0.05/MP (image), $0.08/sec (video); Kling 2.1: $0.10/sec; HunyuanVideo 1.5: $0.08/sec | Flat per-second / per-megapixel billing, no tier-based markup | ❌ Proxy required (AWS US-East deployment) | View Details → | |
| #13 | Helicone | 通过 AI Gateway 代理 100+ Provider 任意模型, AI Gateway + LLM Observability 集成 | 免费层: 10,000 请求/月 + 1GB 存储;Pro $79/月(团队版含告警、HQL 查询);Team $799/月(SOC-2/HIPAA) | 按月订阅 + 用量计费(超过免费额度) | ✅ Open-source self-host + SaaS | View Details → | |
| #14 | FreeModel | DeepSeek V3 / R1, Qwen 2.5 / QwQ-32B | 模型相关(平台加价待确认) | 模型相关 | ✅ Available in China | Credits on registration (amount TBD) | View Details → |
| #14 | Cerebras | Cerebras Llama 3.3 70B, Cerebras Llama 3.1 405B | $0.10/M to $0.60/M tokens(所有模型统一 $0.60/M input) | $0.10/M to $0.60/M tokens(统一 $0.60/M output) | ❌ Proxy required | View Details → | |
| #14 | AI21 Labs (Jamba) | Jamba 1.5 Mini, Jamba 1.5 Large | Jamba 1.5 Mini: $0.20/M tokens, Jamba 1.5 Large: $2/M tokens | Jamba 1.5 Mini: $0.20/M tokens, Jamba 1.5 Large: $2/M tokens | ❌ Proxy required | View Details → | |
| #14 | CoreWeave | NVIDIA HGX H100 ($49.24/hr), HGX H200 ($50.44/hr), HGX B200 ($68.80/hr), A100 80GB ($21.60/hr), L40S ($18.00/hr), L40 ($10.00/hr), GB200 NVL72 ($42.00/hr) — On-Demand/Spot GPU instances, Serverless Inference (W&B Inference, pay-per-token): GLM 5.2, Kimi K2.6/K2.7, DeepSeek R1/V3, Llama 3.x, Qwen — OpenAI-compatible API | GPU On-Demand hourly: NVIDIA HGX H100 $49.24, HGX H200 $50.44, HGX B200 $68.80, A100 80GB $21.60, L40S $18.00, L40 $10.00, GB200 NVL72 $42.00. Spot discounts: H100 $19.71, H200 $20.93, B200 $34.11, A100 $9.65. Serverless Inference is pay-per-token via W&B Inference; Dedicated Inference is GPU-hour / node-based. | ❌ No mainland China direct endpoint. Data centers concentrated in the US; expanded to Indonesia (first APAC region), UK (two operational data centers), and Sweden in 2026. China access requires proxy; cross-Pacific latency is high. | No permanent free tier. Serverless Inference bills per token; GPU instances bill hourly (On-Demand) with spot discounts. The 0EM (Zero Egress Migration) program waives egress fees during migration into CoreWeave. | View Details → | |
| #15 | APIKEY.FUN | GPT-4o, Claude 3.5 Sonnet | 主流模型,价格约为官方 API 的 60-80% | 主流模型,价格约为官方 API 的 60-80% | ✅ Direct access in China | View Details → | |
| #15 | DeepInfra | Meta-Llama-3.3-70B-Instruct, Meta-Llama-3.1-405B-Instruct | Llama 3.3 70B: $0.49/M, DeepSeek-V3: $0.99/M, DeepSeek-R1: $1.09/M | Llama 3.3 70B: $0.73/M, DeepSeek-V3: $0.99/M, DeepSeek-R1: $1.09/M | ❌ Proxy required | View Details → | |
| #15 | NVIDIA NIM | Nemotron-3 Ultra 550B (a55B), Nemotron-3 Super 120B (a12B) | Partner-dependent (Nemotron-3 Ultra 550B: $0.50-0.90/1M tokens) | Partner-dependent (Nemotron-3 Ultra 550B: $1.70-3.60/1M tokens) | ❌ Proxy required (US-based platform, export control restrictions apply) | Free serverless API endpoints for prototyping (no credit card required, generate API key to start) | View Details → |
| #15 | Amazon Bedrock | Claude Opus 4.8 / 4.7 / 4.6 / 4.5, Claude Sonnet 4.6 / 4.5 / 4 | Claude Sonnet 4.5: $6/M, DeepSeek V3.2: $0.62/M, Mistral Large 3: $0.50/M, GPT-5.5: $5.50/M, GPT-5.4: $2.75/M, Grok 4.3: $1.25/M, Qwen3 235B: $0.23/M, Kimi K2.5: $0.60/M | Claude Sonnet 4.5: $30/M, DeepSeek V3.2: $1.85/M, Mistral Large 3: $1.50/M, GPT-5.5: $33/M, GPT-5.4: $16.50/M, Grok 4.3: $2.50/M, Qwen3 235B: $0.91/M, Kimi K2.5: $3/M | ❌ Proxy/VPN required for mainland China users (Bedrock not available in AWS China regions) | AWS Free Tier offers up to $200 credits for new customers (valid 6 months). No permanent free tier for Bedrock. | View Details → |
| #15 | Hyperbolic | DeepSeek R1, DeepSeek R1-0528 | Per-token inference scaled from market GPU rates; H100 from $2.89/GPU-hour, H200 $3.49/GPU-hour | Pay-as-you-go GPU compute billed per hour; reserved discounts for committed capacity | ❌ Proxy required | No permanent free tier and no long-term lock-in. On-demand GPUs are pay-as-you-go billed per hour, with no quota limits and no minimum commitment to start. | View Details → |
| #15 | Fireworks AI | DeepSeek V4 Pro — $1.74 input / $0.145 cached / $3.48 output per 1M tokens (Serverless Standard; Priority $2.61 / $0.218 / $5.22), DeepSeek V4 Flash (0731) — $0.22 input / $0.007 cached / $0.66 output per 1M (Standard) | Serverless (Standard, $/1M tokens): DeepSeek V4 Pro $1.74、V4 Flash $0.22、Kimi K3 $3.00、K2.7 Code $0.95、MiniMax M3 $0.30、Qwen 3.7 Plus $0.40、GPT OSS 120B $0.15、GLM 5.1 $1.40。嵌入按模型档 $0.008-$0.016/1M。按需 GPU 按小时:H100 $7.00、H200 $7.00、B200 $10.00、B300 $12.00、GB300 $18.00(9 月 1 日上调)。 | Serverless 输出($/1M):DeepSeek V4 Pro $3.48、V4 Flash $0.66、Kimi K3 $15.00、MiniMax M3 $1.20、Qwen 3.7 Plus $1.60、GPT OSS 120B $0.60、GLM 5.1 $4.40。缓存输入显著更低(如 DeepSeek V4 Pro $0.145)。训练:SFT 每 1M 训练 token $0.50(≤16B)到 $10.00(>300B);Serverless Training API 仅按 token 计费、无闲置 GPU 成本。 | ❌ No mainland China direct endpoint. As a US-based platform, direct access from the mainland requires proxy or relay; cross-Pacific first-byte latency typically 150-250ms. Closest APAC endpoint is the Tokyo multi-region. | $1 in free serverless credits on signup (pay-per-token, postpaid billing); no permanent free tier. On-demand GPUs and training bill by usage. | View Details → |
| #16 | Replicate | Llama 3.3 70B, DeepSeek-R1 | 模型调用按运行时长 + GPU 型号计费,$0.00025/s/A100起步 | 按运行时长计费(非 token 计费模式) | ❌ Proxy required | View Details → | |
| #16 | SambaNova | Llama 3.3 70B, Llama 3.1 405B | Llama 3.3 70B: $0.50/M, DeepSeek-R1: $0.80/M, Qwen 72B: $0.40/M | Llama 3.3 70B: $0.50/M, DeepSeek-R1: $0.80/M, Qwen 72B: $0.40/M | ❌ Proxy required | View Details → | |
| #16 | 01.AI Yi (零一万物) | Yi-Lightning (智能路由 → DeepSeek-V3/Qwen3-30B-A3B/Yi-Lightning), Yi-Vision-v2 (视觉理解,路由 → Qwen2.5-VL-72B/Yi-Vision-V2) | ¥0.99/1M total tokens(Yi-Lightning,input+output 合并计费) | ¥0.99/1M total tokens(合并计费,不区分 input/output) | ✅ Direct access in China (Chinese company, supports domestic registration + Alipay/WeChat Pay) | Free rate limit tier on registration (limited RPM/TPM), no free initial credits | View Details → |
| #16 | Nebius | DeepSeek V4 Flash, DeepSeek V4 Pro | Per-token OpenAI-compatible inference: DeepSeek-V4-Flash $0.14/M input, $0.28/M output; Kimi K3 $3/$15; MiniMax M3 $0.30/$1.20; Llama 3.3 70B $0.13/$0.40 | Two flavors per model - Base and Fast (-fast suffix) - same outputs, Fast trades higher token price for lower latency via speculative decoding | ❌ Proxy required | No setup fee and no monthly subscription. Token Factory is pay-as-you-go per token, with elastic dynamic rate limits that auto-scale up to 20x your base allocation as sustained usage grows - no capacity reservation needed to start. | View Details → |
| #17 | Anyscale | Llama 3.3 70B, Llama 3.1 405B | Llama 3.3 70B: $0.39/M, DeepSeek-R1: $1/M, Mixtral 8x22B: $0.90/M | Llama 3.3 70B: $0.39/M, DeepSeek-R1: $1/M, Mixtral 8x22B: $0.90/M | ❌ Proxy required | New users get $100 in one-time free credits; small project starter credits ($3-$5) are also available. | View Details → |
| #17 | DigitalOcean Gradient | Llama 3.3 70B, Llama 3.1 405B | Llama 3.3 70B: $0.50/M, DeepSeek-R1: $0.80/M | Llama 3.3 70B: $0.50/M, DeepSeek-R1: $0.80/M | ❌ Proxy required | New accounts get $200 in credits valid for 60 days. There is no permanent free tier; after the credit window, serverless inference bills per token (starting around $0.20 per 1M for the smallest hosted model) and GPU Droplets bill by the hour. | View Details → |
| #17 | MiniMax (Hailuo AI) | MiniMax-01 (Lightning / Turbo / Pro), MiniMax-Text-01 | MiniMax-01 Lightning: ¥0.2/M tokens, Turbo: ¥0.8/M, Pro: ¥2/M, M3: ¥1.5/M | MiniMax-01 Lightning: ¥0.2/M tokens, Turbo: ¥0.8/M, Pro: ¥2/M, M3: ¥1.5/M | ✅ Direct access in China | View Details → | |
| #18 | Stability AI | Stable Diffusion 3.5, Stable Diffusion XL | SD3.5: $0.0065/image, SDXL: $0.01/image, Stable Code: $0.50/M tokens | 按图像/视频/音频输出计费,非 token 计费 | ❌ Proxy required | View Details → | |
| #18 | Novita AI | Llama 3.3 70B, Llama 3.1 405B | Llama 3.3 70B: $0.59/M, DeepSeek-V3: $1.00/M | Llama 3.3 70B: $0.79/M, DeepSeek-V3: $1.00/M | ⚠️ Partial (Singapore node, acceptable latency from China) | View Details → | |
| #18 | Baseten | Llama 3.3 70B, Llama 3.1 405B | Per-second GPU billing (H100: $1.89/hr, A100: $0.83/hr, L40S: $0.54/hr) | Model-dependent (inference time × GPU rate) | ⚠️ Partially available (stable proxy recommended in China) | $30 in inference credits on signup (valid 30 days); custom model hosting free tier: 1 deployed model | View Details → |
| #18 | Mem0 | OpenAI GPT-4o / GPT-4-Turbo / GPT-3.5, Anthropic Claude 3.5 Sonnet / Haiku / Opus | Hobby 免费:10,000 memories + 1,000 retrieval API calls/月;Starter $19/月:50,000 memories + 5,000 calls | Pro $249/月:Unlimited memories + 50,000 retrieval API calls + Graph Memory + 多项目;Enterprise 合同制 | ✅ Open-source (Apache-2.0, 61k+ stars) + Cloud SaaS; self-host works in China without proxy | Hobby free tier: 10,000 memories + 1,000 retrieval API calls/month (unlimited end users), community support | View Details → |
| #18 | Tavily | Search (real-time web + AI summary, 1 credit/call), Extract (clean markdown from URL, 1 credit/page) | Free 1,000 credits/month(滚动 30 天窗口,无信用卡);Pay-as-you-go $0.008/credit(基础 Search 1 credit);Project 起步 $30/month 含 4,000 credits,$0.008/credit 超额;Growth $200/month 含 40,000 credits,$0.006/credit 超额;Pro $1,000/month 含 250,000 credits,$0.005/credit 超额;Enterprise 合同制(custom 配额 + SLA + 无日志合约) | 按 endpoint 计费:Search 1 credit/次;Extract 1 credit/页;Crawl 1 credit/页;Map 1 credit/页;Research 5 credits/次(替代 5-10 Search + 2-3 Extract + 1 合成 LLM);AI Extraction 1 credit/次 | ⚠️ Tavily API hosted on AWS US-East / EU-West; direct access from mainland China 200-400ms latency; recommended Cloudflare Worker or Tencent Cloud Edge Function proxy (50-100ms); no CN node planned; no domestic payment channels | Free forever: 1,000 credits/month (rolling 30-day window), no credit card, no feature gating; Search/Extract/Crawl/Map/AI Extraction each have separate quotas (1,000 Search + 500 Crawl); overage returns HTTP 429, no surprise billing | View Details → |
| #19 | Lepton AI | Llama 3.3 70B, Llama 3.1 405B | Llama 3.3 70B: $0.80/M; Llama 3.1 405B: $3.50/M; Qwen 2.5 72B: $0.80/M; DeepSeek-R1: $2.00/M | Same rate as input for most models (symmetric pricing) | ⚠️ Partially available (ap-northeast-1 Tokyo region is closest; China access requires proxy) | $5 in credits valid for 30 days (enough to fully evaluate any flagship model) | View Details → |
| #19 | Kling AI (Kuaishou) | Kling 3.0 (native 4K video, multi-shot sequencing, Native Audio), Kling 3.0 Omni (multimodal video, video/image reference input) | Per-second video billing. Kling 3.0 Turbo: 720p $0.112/s, 1080p $0.14/s. Kling 3.0: 720p $0.084-0.126/s, 1080p $0.112-0.168/s (Native Audio), 4K $0.42/s. Kling 3.0 Omni: 720p $0.084-0.126/s, 1080p $0.112-0.168/s, 4K $0.42/s, varies by video/audio input. | Image API: Kling Image 3.0 $0.028/image (1K/2K), 3.0-omni $0.028-0.056/image (up to 4K), Image 2.1 $0.014-0.028/image, Image O1 $0.028/image, Multi-Shot $0.07/call. Video list price 1 Unit = $0.14; image list price 1 Unit = $0.0035. | ✅ Official Chinese service available (Kuaishou first-party, mainland cloud nodes available via Kling Open Platform); the global klingai.com developer API can require network setup depending on region | New users get a small one-time free credit grant on the platform to try video/image generation; there is no permanent free API tier — API usage is billed per generated second (video) or per image. | View Details → |
| #20 | Hugging Face | Llama 3.3 70B, Llama 3.1 405B | 推理 API: $0.04-0.60/M tokens 取决于模型 | 推理 API: $0.04-0.60/M tokens | ❌ Proxy required (huggingface.co blocked in China) | View Details → | |
| #20 | ElevenLabs | Eleven Multilingual v2, Eleven Turbo v2.5 | TTS: $0.03-0.30/1K chars depending on model/quality; STT: $0.47/hour; Sound Effects: $0.04/request | Same as input rate per model | ⚠️ Partially available (stable proxy recommended in China) | View Details → | |
| #20 | Ideogram | Ideogram 3.0, Ideogram Turbo | Ideogram 3.0: $0.04/image (standard), $0.08/image (Turbo) | 4 images per generation (default), additional images via paid credits | ❌ Proxy required | View Details → | |
| #20 | Lambda | NVIDIA HGX B200 180GB (on-demand GPU instances; $6.69/GPU/hr), NVIDIA H100 SXM 80GB ($3.99/GPU/hr on-demand) | GPU instance on-demand per GPU/hr: NVIDIA B200 SXM6 $6.69, H100 SXM $3.99, A100 SXM 80GB $2.79, A100 SXM 40GB $1.99, GH200 $2.29, A100 PCIe $1.99, A6000 $1.09, A10 $1.29, V100 $0.79. 1-Click Cluster reserved per GPU/hr: B200 $9.86 (16 GPU) to $8.87 (256+), H100 $6.16 to $5.54. Superclusters / Private Cloud (4000+ GPUs) via sales. No egress fees. | Billed by the second per GPU instance (no idle GPU cost); clusters and reserved capacity carry a minimum commitment (2 weeks to 1 year). Managed orchestration: Managed Kubernetes and Slurm included at no extra per-node fee; Lambda Stack one-line install. No per-token inference pricing — Lambda's Inference API is winding down; self-hosted inference billed as GPU runtime. | ❌ No mainland China direct endpoint; data centers in the US (California and other regions), China access requires proxy or overseas relay. | No permanent free GPU tier; usage-based billing after signup (GPU instances billed per second). 1-Click Clusters and Superclusters require prepaid / committed terms (2 weeks to 1 year). | View Details → |
| #21 | Deepgram | Nova-3 (STT, Monolingual), Nova-3 (STT, Multilingual) | STT: $0.0043-0.012/min; TTS: $0.015-0.030/1K chars | N/A | ⚠️ Partially available (proxy required, no mainland China data center) | View Details → | |
| #21 | RunPod | B300 288GB HBM3e, B200 180GB | Per-second GPU: H100 SXM $2.99/hr, A100 SXM $1.49/hr, B200 $5.89/hr, L40S $0.99/hr, RTX 4090 $0.69/hr | Same as input (per-GPU-second — no idle fees, no egress fees) | ⚠️ Partially available (31 global regions; CN direct access unstable — proxy recommended) | $5 starter credit on signup, no credit card required; per-second billing means spiky workloads never pay for idle | View Details → |
| #21 | Luma AI | Ray 3.2 (flagship video, 1080p, multi-keyframe up to 16 frames, V2V up to 20s), Ray 3.14 (previous-gen video model) | Credit-based subscription. Plans: Plus $30/mo (10,000 credits), Pro $90/mo (40,000 credits), Ultra $300/mo (150,000 credits). Ray 3.2 video: 1080p text-to-video 400 credits/5s, 720p 100 credits/5s, Draft 20 credits/5s; Seedance 2.0 1080p 240 credits/sec, 4K 959 credits/sec. | Image credits: Uni-1 30 credits/image, Seedream 1-3 credits/image, GPT Image 2 from 3 (Low-1K) to 255 (High-4K) credits. Audio: ElevenLabs v3 TTS 21 credits/1,000 chars. Utilities: background removal 1 credit/image, reframe video 32 credits/sec, upscale to 4K 17 credits/sec. | ❌ Proxy required (no mainland China direct access; API served from global AWS/GCP edge regions) | No permanent free API tier. New users get a small one-time free credit grant on the platform, but sustained API usage requires at least the Plus plan at $30/mo (10,000 credits). | View Details → |
| #22 | Chroma | All-MiniLM-L6-v2 (default, 384-dim, SBERT), All-mpnet-base-v2 (768-dim, higher quality SBERT) | Chroma OSS 完全免费(Apache 2.0),本地/嵌入式运行;Chroma Cloud Free 永久层:50K 向量、5K 查询/月、0.5 GB 存储;Pro $0.30/M 向量-月 + $0.01/M 查询(月最低 $10);Enterprise 合同制(自定义节点、HIPAA、SOC 2) | 存储(GB-月)+ 向量数 + 查询数 三维度计费;无 per-embedding API 费;Default Embedding Function 本地推理零成本;BYO Embedding 由第三方 API 计费 | ✅ Chroma OSS runs fully local (on any server, zero latency); Chroma Cloud on AWS us-east-1 / eu-west-1 / ap-southeast-2 (Singapore node planned H2 2026), ~200-400ms latency from mainland China | Chroma OSS permanently free, Apache 2.0, no feature limits, no vector count cap (limited by local hardware); Chroma Cloud Free forever: 50K vectors + 5K queries/mo + 0.5 GB storage + 1 Collection + community support | View Details → |
| #22 | Voyage AI | voyage-4-large, voyage-4 | voyage-4-large: $0.12/M, voyage-4: $0.06/M, voyage-4-lite: $0.02/M, voyage-context-3: $0.18/M, voyage-code-3: $0.18/M, voyage-multimodal-3.5: $0.12/M | (embedding API — no output tokens; rerank-2.5: $0.05/M tokens, rerank-2.5-lite: $0.02/M tokens) | ⚠️ Partial (site accessible from China, API stability in mainland China unverified) | 200M free tokens per account (text embedding + reranker), 50M free tokens for legacy domain models | View Details → |
| #22 | Runway | Gen-4 Turbo, Gen-4 Alpha | Gen-4 Turbo: 50 credits/5s clip (~$0.50), Gen-4 Alpha: 100 credits/5s (~$1.00), Act-Two: 150 credits/5s (~$1.50), Frames: 200 credits/sequence (~$2.00) | Credit-based system (~$0.01/credit). Credit packs: $12 (1,000 credits) to $200 (25,000 credits). Enterprise: custom pricing. | ❌ Proxy required (AWS US-East + GCP US-Central hosting) | No free API tier. Minimum $12 for 1,000 credits (~20 Gen-4 Turbo clips). Free web app available but no API access. | View Details → |
| #22 | TwelveLabs | Pegasus 1.5 (flagship multimodal video understanding: frames + audio + speech + on-screen text), Pegasus 1.2 (previous-gen, lower-cost option: frames + audio + speech) | Free plan: 600 minutes video indexing + API usage (Pegasus 1.2 $0.042/min indexing one-time, $0.021/min input video, $0.0075/1k tokens output text, $0.0015/min monthly embedding infrastructure). Developer plan: paid overage, 90-day index retention lifted to unlimited. | Same per-minute and per-1k-tokens billing. Marengo embeddings billed separately by indexed minute. | ❌ Proxy required (no mainland China direct endpoint; API hosted on AWS us-west-2 / us-east-1; SDKs available in Python / Node / Go / Java) | Free plan: 600 minutes of free video indexing on signup (cumulative, no credit card required). Full API surface available (Pegasus 1.2 / 1.5, Marengo 3.0, Search, Embed, Analyze & Segment). Index retention is 90 days on Free; upgrading to Developer lifts retention to unlimited. | View Details → |
| #22 | Mixedbread AI | Toast 1 — specialized search model (2026-08-13): matches/outperforms Claude Opus 5 and GPT-5.6 Sol on knowledge work, up to 10x cheaper and 12x faster; input $0.50 / cached $0.06 / output $1.20 per 1M LLM tokens (launch $0.30 / $0.036 / $0.72), mxbai-embed-large-v1 — 1,024-dimension open embedding model (Apache-2.0), 50M+ downloads across the mxbai family | 按量计费,三块:索引(Fast $1.50/百万内容 token;High Quality 含 OCR/转写/摘要/多模态增强 $3/百万)、搜索(Grep $0.10/千次、Semantic $4/千次、Toast 1 $1/千次,重排 +$3.50 或 +$1.50/千次)、存储 $0.50/百万内容 token/月。Toast 1 按 LLM token:输入 $0.50/百万(发布价 $0.30)。 | Toast 1 输出 $1.20/百万 LLM token(发布价 $0.72,40% 折扣);缓存输入 $0.06/百万(发布价 $0.036),缓存写入免费。Semantic 检索重排加价 $3.50/千次。 | ❌ No mainland China direct endpoint. Mixedbread is a US/Europe-based platform (legal entity mixedbread ai inc., Berlin/San Francisco), so direct access from the mainland requires a proxy or relay, with cross-Pacific first-byte latency typically 150-250ms. | Starter plan grants $5 in one-time credits with no card required; 3 workspace users, 10 stores, 100 requests/minute — enough to evaluate a real corpus. No permanent free tier. | View Details → |
| #23 | Crusoe | DeepSeek V3 0324 (legacy frontier: $0.50 input / $1.50 output per 1M tokens), DeepSeek V4 Flash (efficient frontier: $0.14 input / $0.28 output per 1M tokens, 70B tier) | Pay-as-you-go per 1M tokens, four tiers by parameter count: <16B ($0.40) / 16B-70B ($2.50) / 70B-300B ($6.00) / >300B ($10.00). Serverless Fine-Tuning follows identical 4-tier structure. Managed Inference spans DeepSeek V3/V4 Pro/V4 Flash, GLM 5.1/5.2, GPT-OSS 20B/120B, Gemma 4 31B-it, Kimi K2.6, Llama 3.1/3.3, Nemotron 3 family (VoiceChat/Ultra/Lightning/Nano/Super/Nano Omni), Qwen3 family (8B/235B/Qwen3.5/Qwen3.6), Yutori n1.5. Cached tokens at $0.03-$1.50 per 1M. | Self-Serve Deployments: NVIDIA H100 80GB HGX $5.50/hr, NVIDIA H200 141GB HGX $6.00/hr (dedicated endpoints for open + fine-tuned models). Tailored Deployments and Provisioned Throughput via sales engagement. Managed Kubernetes $0.10/cluster hour; Container Registry $0.10/GiB-month; Object Storage $0.06/GiB-month. | ❌ No mainland China direct endpoint. Managed Inference API hosted at Crusoe's US data centers (Colorado, Texas, and other US locations); China access requires proxy. | No permanent free tier; pay-as-you-go only. Volume discounts via sales contact. Serverless Fine-Tuning starts at $0.40/M tokens for sub-16B models. | View Details → |
| #23 | Letta | Letta Agents SDK — open-source stateful agent runtime (Apache-2.0), with first-class persistent memory (core/archival/recall blocks) and runtime memory editing, Letta Code (`@letta-ai/letta-code`, Node.js 22.19+) — memory-first CLI coding agent, #1 on Terminal-Bench (Dec 2025 launch) | 按用量:Free $0(最多 3 个有状态 agent、BYOK)、Pro $20/月(最多 20 个 agent、Letta Auto 周/月额度 + 按量超额)、Teams Pro 按席位(共享 agent、权限)、Developer Plan 纯 API 调用按信用(无 agent 上限)、Enterprise 商务报价(自托管/定制模型)。服务端工具按 CPU 时间 $0.00015/秒计费。 | Letta Auto 用量超出 Pro 套餐限额后按 pay-as-you-go 模型价(参考 platform.letta.com/models)。Mods、Skills、Conversations 等功能均包含在套餐内,不单独计费。远程 MCP 工具由 MCP 供应商执行,不消耗 Letta 信用。 | ❌ No mainland China direct endpoint. Letta, Inc. is registered in San Francisco and platform.letta.com is a single global endpoint, so production traffic from mainland China requires a proxy or relay (typical cross-Pacific first-byte latency ~150-250ms). The agent harness is fully open-source (Apache-2.0), so self-hosting on your own infrastructure is a viable workaround. | Free $0/mo with up to 3 stateful agents and BYOK; Letta Auto and Skills are throttled — usable for evaluation and small personal-assistant use cases. | View Details → |
| #24 | Weaviate | Snowflake Arctic-embed-m-v1.5, Snowflake Arctic-embed-m-v2.0 | Free 永久层:100,000 对象 + 1 集合 + 7 天备份;Flex 月付 $45 起 + 从 $0.00465/1M 维度 + $0.12/GiB 存储;Premium 预付从 $400/月 | Dedicated 合同制 from $400/月 起;Vector dimension、Storage、Backup 三维度独立计费;Embeddings 按 token 用量计费 | ⚠️ Servers in AWS/GCP regions, ~200-400ms latency from China | Free forever: 100,000 objects, 1 collection, 7-day backup; 2,000 Embedding requests/day; Query Agent 1,000 req/mo | View Details → |
| #25 | Writer | Palmyra X5, Palmyra X4 | Palmyra X5: $0.60/M, Palmyra X4: $2.50/M | Palmyra X5: $6.00/M, Palmyra X4: $10.00/M | ⚠️ Partial (site accessible from China, API stability unverified) | 14-day free trial, no credit card required (Starter plan) | View Details → |
| #25 | Reka | Reka Core — top-tier multimodal model (text + image + audio + video); input $2.00 / output $6.00 per 1M tokens, image $0.02 / video $0.08 / audio $0.02 per minute, Reka Flash — cost-efficient model for everyday tasks (text + image + audio + video); input $0.80 / output $2.00 per 1M tokens, image $0.01 / video $0.06 / audio $0.015 per minute | Reka Edge $0.10/M 输入;Reka Flash $0.80/M;Reka Core $2.00/M;Reka Flash Research 按请求计费($25-$60/千次)。按量付费,无最低承诺。 | Reka Edge $0.10/M;Reka Flash $2.00/M;Reka Core $6.00/M。多模态(图像、视频、音频)单独按"每分钟"计费,Flash 模型视频 $0.06/分钟、Core $0.08/分钟;图像按调用计费(Flash $0.01、Core $0.02)。 | ❌ No mainland China direct endpoint. Reka is registered at 530 Lawrence Expressway, Sunnyvale, CA; api.reka.ai is a single global endpoint, so production traffic from mainland China requires a proxy or relay (typical cross-Pacific first-byte latency ~150-250ms). The open weights can be self-hosted on NVIDIA GPUs with ≥24GB VRAM via vLLM to bypass direct-access limits. | Sign up at platform.reka.ai to obtain an API key; billing is purely pay-as-you-go with no minimum commitment, so teams can prototype with a small top-up. No permanent free tier or monthly free credit. | View Details → |
| #26 | Qdrant | FastEmbed: BAAI/bge-small-en-v1.5, FastEmbed: BAAI/bge-base-en-v1.5 | Free 永久层:1 GB 存储 + 0.5M 向量,无限期;Cloud Standard 月付 $25 起(预付 $250/yr)按存储 $0.06/GB-月;Cloud Pro 月付 $80 起 按存储 $0.04/GB-月 + 高 IOPS;Dedicated 合同制 $2,500/月起 | 按存储维度(GB)与可选性能包(Power Tiers)计费;无 API 调用费、无 per-vector 嵌入费;FastEmbed 按 token 数本地推理;BYO Embedding 由第三方 API 计费 | ⚠️ Cloud runs on AWS Frankfurt / N. Virginia / Sydney regions, ~200-400ms latency from mainland China; OSS can be self-hosted in Tencent/Aliyun CN regions | Free forever: 1 GB storage + 0.5M 1015-dim vectors; 2 CPU/0.5 GiB RAM instance; unlimited API requests; community Discord support | View Details → |
| #27 | fal.ai | FLUX.2 Pro (text-to-image, $0.03/MP first MP + $0.015/extra MP), FLUX1.1 [pro] (ultra-fast text-to-image, per-megapixel) | 按请求/按时长计费,无月费。图像生成:FLUX.2 Pro $0.03/第一个 MP + $0.015/额外 MP、FLUX1.1 [pro] 按 MP、Nano Banana 2 $0.08/图、FLUX.1 [schnell] 无定价显示。视频:Kling v3 $0.084-0.168/秒、Seedance 2.0 $0.3034/秒(720p)。TTS:通过 ElevenLabs/MiniMax 等第三方提供。余额模式:预付 $10+ 存入信用余额,用完为止。开源模型按秒计费,秒级颗粒度。 | ✅ Available from Mainland China (fal.ai uses global CDN, no proxy needed); latency varies by model inference node, typically 200-500ms | No free API tier (prepay $10+ for credit balance); no monthly fees, no subscriptions, no hidden fees; serverless auto-scales to zero when idle | View Details → | |
| #28 | Aleph Alpha | Pharia-1-LLM-7B-control, Pharia-1-LLM-7B-control-aligned | Enterprise contract — no public per-token price (quote-based; PhariaAI on-prem deployment custom-priced) | Same — enterprise quote, no public per-token list | ❌ Not applicable (German sovereign-AI vendor, subject to EU export controls) | No public free tier — self-serve signup creates an account, but pricing is enterprise contract only | View Details → |
| #28 | Cartesia | Sonic (real-time TTS), Sonic 2 (multilingual, voice cloning) | Sonic: $0.03/1K chars (Standard), $0.06/1K chars (Pro), $0.12/1K chars (Turbo); Sonic 2: $0.04-0.15/1K chars | TTS character-based pricing (output = audio), per-second for streaming | ⚠️ Partially available (stable proxy recommended in China) | 10,000 characters/month free tier (≈10 minutes), no credit card required | View Details → |
| #28 | Hume AI | EVI 3 (Empathic Voice Interface 3), OCTAVE TTS (voice design + cloning) | EVI 3: ~$0.096/min (pay-per-second); OCTAVE TTS: $0.048/1K chars; Expression Measurement: $0.0008/sec | Same as input (most endpoints symmetric) | Proxy required (US export controls) | Free credits on signup (amount per account), enough to trial EVI 3 | View Details → |
| #28 | Modal | B300 / B200 / H200 / H100 / A100 / L40S / A10 / L4 / T4 (GPU), Self-hosted LLM inference (vLLM, SGLang, TensorRT-LLM) | 按 GPU 秒计费:H100 $0.001097/s、A100 80GB $0.000694/s、T4 $0.000164/s | 同输入(按 GPU 时间,无 idle 费用) | ❌ Proxy required (infrastructure on AWS/GCP overseas regions; direct CN access unreliable) | View Details → | |
| #30 | AssemblyAI | Universal-3.5 Pro (Async, 18 languages, native code switching, best diarization), Universal-2 (Async, 99 languages, balanced accuracy) | 按音频小时计费:Universal-3.5 Pro $0.21/hr、Universal-2 $0.15/hr(pre-recorded) | N/A(speech-to-text 计费基于音频时长,无 token 概念) | ❌ Proxy required (US-based infrastructure, CN direct access unreliable) | View Details → | |
| #30 | Pinecone | Serverless (Standard / Enterprise), Pod-based (s1 / p1 / p2) | Serverless Standard: $50/月最低,$4-$4.50/M读单元,$16-$18/M写单元,$0.0005/ingestion单元 | Serverless Enterprise: $500/月最低,$6-$6.75/M读,$24-$27/M写,$0.001/ingestion单元(multi-modal);Pod-based按节点小时计费 | ⚠️ SaaS-only — no open-source self-host option (proprietary managed service) | View Details → | |
| #39 | Together AI | DeepSeek V4 Pro, Qwen3.6-Plus / Qwen3.7-Max | $0.03-$1.74/MTok(模型相关) | $0.12-$4.50/MTok(模型相关) | ❌ Proxy required (US-based service) | $1 initial credit on signup (requires payment method) | View Details → |
| # | FriendliAI | GLM-5.2, GLM-5.1 | GLM-5.2/5.1: $1.4/M, DeepSeek-V3.2: $0.5/M, Qwen3-235B: $0.2/M, MiniMax-M2.5: $0.3/M, Gemma-4-31B: $0.14/M | GLM-5.2/5.1: $4.4/M, DeepSeek-V3.2: $1.5/M, Qwen3-235B: $0.8/M, MiniMax-M2.5: $1.2/M, Gemma-4-31B: $0.4/M | ⚠️ US + Korea servers, ~200-400ms latency from China | No free tier — per-token billing, no minimum | View Details → |
| #40 | Suno | v5.5 (latest model, royalty-free music), v5 (premium model) | Free: $0 (10 songs/day, non-commercial); Pro: $10/mo (500 songs, commercial rights); Premier: $30/mo (2,000 songs, Studio) | Annual: Pro $96/yr (20% off), Premier $288/yr (20% off); credit top-ups available as add-ons | ⚠️ Limited access from China (stable overseas proxy required), no official China endpoints | Free tier: 10 songs per day (non-commercial use only), no credit card required, available immediately on signup | View Details → |
| #40 | Liquid AI | LFM2.5-8B-A1B (MoE, 8B total / 1.5B active, 128K context) — open weights, free download / run / fine-tune under the LFM Open License (royalty-free until $10M annual revenue), LFM2.5-2.6B (dense, 128K context, agentic tool calling) — DSpark draft variant up to 3.18× GPU / 2.87× on-device decode speedup | 无按 token 云 API 定价——LFM 权重免费下载/运行/微调(LFM Open License,公司年收入超过 1000 万美元前免版税)。推理要么在自有硬件上自托管,要么通过 LEAP SDK 本地部署;无 per-token 请求费用。 | 无 per-token 输出计费。模型在 CPU/GPU/NPU 上本地或自托管运行,适合对延迟与隐私敏感的高频或离线工作负载(Liquid 的定位:「为何为云 API 按 token 付费,而非本地推理」)。 | ⚠️ No managed cloud API, so there is no China direct-connect issue. Model weights download freely and run locally or self-hosted, including inside mainland China (data-sovereignty friendly). No cloud API gateway or proxy needed; token processing happens entirely on local hardware. | All 14+ LFM models are free to download, run, and fine-tune under the LFM Open License — including commercially (until your company exceeds $10M annual revenue). Research, education, and non-profit use is always free with no revenue limit. | View Details → |
| # | Nscale | Moonshot AI Kimi K2.5 — open chat model via serverless inference; per-1M input/output token pricing (verified on nscale.com/product/inference featured model list), Alibaba Cloud Qwen3 8B / Qwen3 4B Thinking / Qwen3 4B Instruct — open chat models with reasoning variants; per-1M token pricing | Nscale 采用充值余额(credit)按量计费,价格按模型分列,在 Nscale 控制台 AI Services → Models 与 /v1/models API 中按每 1M 输入 token 报价(文本模型);图像模型输入按 0 计费。定价结构经 serverless-openapi.yaml ModelPricing schema 验证(input = 每 1M 输入 token 美元价)。 | 输出按每 1M 输出 token 计费(文本模型);图像模型按每百万像素计费(ModelPricing.output 定义为 per-million output pixels)。具体美元单价在控制台与 /v1/models 端点内按模型展示,注册充值后可查。 | ❌ No mainland China direct endpoint. inference.api.nscale.com is a single global endpoint; data centers concentrate in Europe (Loughton UK, Glomfjord/Narvik Norway) and the US (Texas, West Virginia), so China production traffic needs a proxy and transatlantic/transpacific first-byte latency is typically 200-350ms. The platform emphasizes EU data sovereignty and in-country processing, which suits European compliance workloads more than China-region low latency. | No permanent free tier. New users sign up via Google SSO; the quickstart notes early users may be eligible for free promotional credits (not guaranteed). By default you must top up credit before calling the serverless inference API. | View Details → |
| #41 | Morph | Kimi K3 (2.8T MoE) — flagship Fable-tier open-weight model at ~100 tok/s with 1M context; per-1M input/output token pricing (verified on docs.morphllm.com llms.txt Fast Models table 2026-08-28), GLM-5.2 (744B MoE) — Opus-tier open-weight model with 1M context; OpenAI `service_tier` parameter supported for default/standby processing | 按模型分别计费,按每 1M 输入 token 计(文本模型)。从 llms.txt 验证(2026-08-28):Kimi K3 $2.90、Qwen 3.5 397B $0.50、GLM-5.2 744B $1.10、GLM-5.3-Flash $0.15、MiniMax M3 $0.30、MiniMax M2.7 $0.279、DeepSeek V4 Flash (beta) $0.12、Qwen 3.8/3.6 27B $0.289、Gemma 4 31B $0.14。Qwen 3.5 397B 缓存输入 $0.30 per 1M(其他模型无独立缓存价)。Fast Apply 用 `auto` 模型按 prompt+completion 两档计(OpenRouter 验证 morph-v3-fast $0.0008/$0.0012 per 1k、morph-v3-large $0.0009/$0.0019 per 1k),Compact、Reflexes、WarpGrep 与 Router 按事件计费(详见 pricingEN)。 | 按模型分别计费,按每 1M 输出 token 计(文本模型)。从 llms.txt 验证:Kimi K3 $14.00、Qwen 3.5 397B $3.50、GLM-5.2 744B $4.10、GLM-5.3-Flash $0.50、MiniMax M3 $1.20、MiniMax M2.7 $1.20、DeepSeek V4 Flash (beta) $0.278、Qwen 3.8/3.6 27B $2.40、Gemma 4 31B $0.40。Fast Apply 输出按 morph-v3-fast $1.20 / morph-v3-large $1.90 per 1M token;Fast Apply 用 `auto` 推荐档(5,000-10,500 tok/s, ~98% 精度)按实际路由模型分档计费。Kimi K3 `service_tier: standby` 走 GLM-5.2 同等待遇,标准按 token 计。 | ❌ No mainland China direct endpoint. Base URL https://api.morphllm.com is a single global endpoint, hosted by AutoInfra, Inc. (YC-backed US company). OpenRouter routes morph/morph-v3-fast and morph/morph-v3-large through multi-region, but the official Morph API is not available from mainland China without a proxy. China production traffic sees typical 200-400 ms transpacific first-byte latency. | No permanent free tier (no zero-cost starter plan). New signups get an API key and are billed pay-as-you-go from the first token. Morph offers up to $5,000 in Startup Credits for qualified startups, applied via a separate sales contact form (morphllm.com home page). The homepage CTA 'Start free. Get API Key' refers to signup friction, not to a free quota. | View Details → |
| #42 | Firecrawl | Scrape (v2 /scrape) — convert any URL into clean markdown / HTML / structured JSON via the JSON-mode schema option; 1 credit per page, Crawl (v2 /crawl) — recursively crawl a website and return content for every linked page; 1 credit per page scraped | 纯 credit-based 按 endpoint 与特性消耗,每月按 plan 重置额度;额外 credit 按 $5 一档购买(每档 1k–5k credits 不等)。从 firecrawl.dev/pricing 与 docs.firecrawl.dev/billing.md 验证(2026-08-29):Free 1k/月 $0;Hobby 5k/月 $19(年付 $16);Standard 100k/月 $83(年付 $99);Growth 500k/月 $333(年付 $399);Scale 1M/月 $599;Enterprise Custom。Credit 单价:Scrape / Crawl / Map 1 cr/page、Search 2 cr/10 results、Monitor 1 cr/page/check、Interact 2–7 cr/browser-minute。Batch Scrape / Extract 按子操作同价。Agent 5 daily runs free,超量动态计费。 | 无独立 output 维度——Firecrawl 输出的是 scraped markdown / HTML / JSON 内容,计费仅按输入端的 credit 消耗与并发浏览器占用。出站数据本身不另收费;唯一例外是 Search 在 ZDR(zero-data-retention)企业档下按企业价另议。 | ❌ No mainland China direct endpoint. Firecrawl is headquartered in San Francisco (Y Combinator W23 batch) with a single-region API base URL at https://api.firecrawl.dev; the billing and rate-limits docs do not list any Asia-Pacific or China routing layer. China production traffic requires a proxy or transit, with typical 200-400 ms transpacific first-byte latency. The Free tier's 1,000 credits work behind a proxy for evaluation, but stable production traffic should either run a self-hosted Firecrawl deployment (Docker / Kubernetes) or sit behind a corporate proxy. | Permanent Free tier: 1,000 credits per month with no credit card required, full access to Scrape / Crawl / Map / Search / Batch Scrape / JSON mode and every other endpoint; 2 concurrent browsers and a 50,000 max queued jobs ceiling. Designed for evaluation and low-volume use; credits reset at the start of each month and do not roll over. Agent's 5-daily-free runs apply across every plan, including Free. | View Details → |