AI API 厂商大全
对比所有主流 AI API 提供商的价格、模型和功能
| # | 厂商 | 热门模型 | 输入 | 输出 | 国内可用 | 免费额度 | |
|---|---|---|---|---|---|---|---|
| #1 | OpenAI | GPT-5, GPT-5-mini | GPT-5: $1.25/M, GPT-5-mini: $0.25/M, GPT-4o: $2.50/M, GPT-4o-mini: $0.15/M | GPT-5: $10/M, GPT-5-mini: $2/M, GPT-4o: $10/M, GPT-4o-mini: $0.60/M | ❌ 需代理(受出口管制限制,国内用户需使用代理访问) | 新账户赠 $5 额度,有效期 3 个月 | 查看详情 → |
| #2 | Anthropic Claude | Claude Sonnet 5, Claude Opus 5 | Sonnet 5 (intro): $2/M, Opus 5: $5/M, Sonnet 5/4.5: $3/M, Opus 4.8: $5/M, Haiku 4.5: $0.80/M | Sonnet 5 (intro): $10/M, Opus 5: $25/M, Sonnet 5/4.5: $15/M, Opus 4.8: $25/M, Haiku 4.5: $4/M | ❌ 需代理(受美国出口管制,国内用户需代理) | Claude 免费版有限额度(网页端);API 端无赠送额度;Sonnet 5 限时价 $2/M(至 2026-08-31) | 查看详情 → |
| #2 | Azure OpenAI(微软 Azure) | GPT-4o, GPT-4o-mini | GPT-4o: $2.50/M, GPT-4o-mini: $0.15/M, o1: $15/M, o3: $10/M | GPT-4o: $10/M, GPT-4o-mini: $0.60/M, o1: $60/M, o3: $40/M | ⚠️ 世纪互联运营版(国内可用,模型比国际版少) | 新用户 $200 免费额度(30天有限) | 查看详情 → |
| #3 | OpenRouter | GPT-4o, Claude Opus/Sonnet/Haiku | $0.065/M to $15/M 取决于模型,平均 $0.50/M | $0.26/M to $75/M 取决于模型,平均 $2/M | ❌ 需代理(未封锁,但国内访问不稳定) | 无初始免费额度,免信用卡注册 | 查看详情 → |
| #4 | DeepSeek(深度求索) | DeepSeek-V3, DeepSeek-R1 | DeepSeek-V3: ¥0.14/M token (≈$0.02), R1: ¥0.14/M | DeepSeek-V3: ¥0.28/M token (≈$0.04), R1: ¥0.28/M | ✅ 国内直连 | 注册送 ¥500 免费额度(约500万tokens) | 查看详情 → |
| #4 | xAI Grok | Grok 4.6, Grok 4.6 Fast | Grok 4.6: $2/M tokens (fast variant $4/M) | Grok 4.6: $6/M tokens (fast variant $12/M) | ❌ 需代理 | 无免费 API 额度(Grok Build 试用;初期免信用卡) | 查看详情 → |
| #4 | Gemini 多模态 | Gemini Omni Flash (multimodal, 1M context, audio+video+text), Nano Banana 2 Lite (image generation, $0.034/1K tokens) | Omni Flash: $0.15/M input tokens (multimodal); Nano Banana 2 Lite: $0.034/1K image tokens; Veo 3.1: $0.10/sec video | Omni Flash: $0.60/M output tokens (multimodal); Nano Banana Pro: $0.12/1K image tokens | ❌ 需代理(Google AI Studio 在中国大陆不可直接访问,需稳定代理) | Free tier: Gemini Omni Flash 15 RPM, Veo 3.1 2 generations/day, Imagen 4 10 images/day — no credit card required | 查看详情 → |
| #5 | Google Gemini | Gemini 2.5 Pro, Gemini 2.5 Flash | 2.5 Pro: $1.25-10/M, 2.5 Flash: $0.15-1.25/M, 2.0 Flash: $0.10/M | 2.5 Pro: $5-40/M, 2.5 Flash: $0.30-5/M, 2.0 Flash: $0.40/M | ❌ 需代理(Google 云服务在国内不可用) | 免费版:2.5 Pro/Flash 有免费 Rate Limit(15-30 RPM),超出按量计费 | 查看详情 → |
| #5 | 通义千问(阿里) | Qwen3.8-Flash-Next, Qwen3.8-Flash | Qwen3.8-Max: ¥8/M, Qwen3.5-Max: ¥4/M, Qwen3.5-Plus: ¥2/M | Qwen3.8-Max: ¥24/M, Qwen3.5-Max: ¥12/M, Qwen3.5-Plus: ¥6/M | ✅ 国内直连 | 新用户 100万 tokens 免费额度(90天有效);Qwen3.5 系列开源可自部署 | 查看详情 → |
| #6 | Perplexity AI | Sonar Pro, Sonar | Sonar Pro: $3/M, Sonar: $1/M | Sonar Pro: $15/M, Sonar: $1/M | ❌ 需代理 | API 无免费额度,Pro 订阅 $20/月(5000次搜索查询) | 查看详情 → |
| #6 | 美团 LongCat | LongCat-2.0 (1.6T MoE), LongCat-2.0 INT4 | OpenRouter: $0.60/M token (48B active MoE) | OpenRouter: $2.40/M token (48B active MoE) | ✅ Domestic self-host + OpenRouter (proxy) | MIT licensed weights on Hugging Face; self-host on 8x H100+ GPU cluster | 查看详情 → |
| #6 | Meta Model API | Muse Spark 1.1 (muse-spark-1.1) | $1.25/1M tokens | $4.25/1M tokens | ❌ 需代理(Meta Model API 目前仅面向美国开发者公开预览,中国大陆需代理访问) | Free tier: 60 RPM, 2M TPM limit; free credits on signup for US developers (public preview) | 查看详情 → |
| #6 | 阶跃星辰(Stepfun) | Step-2 (1.2T multimodal flagship, text+image+audio+video, 128K context), Step-2-mini (200B multimodal, faster Step-2 variant) | Step-2: ¥6/M, Step-2-mini: ¥1/M, Step-1: ¥4/M, Step-R: ¥8/M | Step-2: ¥18/M, Step-2-mini: ¥3/M, Step-1: ¥12/M, Step-R: ¥24/M | ✅ 国内直连(platform.stepfun.com 端点,国内BGP优化) | 新用户 100万 tokens 免费额度(注册后30天有效) | 查看详情 → |
| #7 | Exa | Search — neural semantic search across the web; 6 search types (auto / instant / fast / deep-lite / deep / deep-reasoning); returns 10 results by default, $1/1k per extra result above 10; AI page summaries $1/1k pages, Contents — full page text / highlights / summaries for known URLs; $1/1k pages per content type (text, highlights, summary billed separately) | Pure pay-as-you-go by endpoint and search-type. Search $7/1k requests (base 10 results) + $1/1k per extra result + $1/1k AI summaries; Contents $1/1k pages; Answer $5/1k; Monitors $15/1k; Deep Search $12-15/1k; Agent fixed effort $0.012-$1.00/request or usage-based $0.10/ACU + tool calls (default $5 auto cap, $20 max cap). New accounts get $20 in free credits (~2,800 searches); Free Tier adds $10/month. No subscription, no minimum spend. | Same as input. Enterprise custom volume + Zero Data Retention + SLA + postpaid invoice. | ❌ Proxy required. Exa primary API runs at https://api.exa.ai from a single-region US deployment (San Francisco). No documented mainland China endpoint as of 2026-08-30. Mainland China production traffic typically needs a proxy or relay; transpacific first-byte latency is usually 100-300 ms. Enterprise plans can negotiate custom routing and dedicated infrastructure for regulated workloads (HIPAA / GDPR data residency). dashboard.exa.ai/onboarding provides agent-friendly onboarding that generates integration snippets tailored to the developer's stack, but API key issuance still requires stable international network access from China. | Permanent $20 signup credit (~2,800 Search calls) + automatic $10 Free Tier monthly credit refill. No credit card, no minimum spend. Free credits apply to every endpoint. | 查看详情 → |
| #7 | 阿里云百炼 | Qwen3.5-Max, Qwen3.5-Plus | Qwen3.5-Max: ¥4/M, Qwen3.5-Plus: ¥2/M, Qwen3.5-72B: ¥4/M, QwQ: ¥2/M | Qwen3.5-Max: ¥12/M, Qwen3.5-Plus: ¥6/M, Qwen3.5-72B: ¥12/M, QwQ: ¥8/M | ✅ 国内直连 | 新用户 100万 tokens 免费额度(90天有效) | 查看详情 → |
| #8 | 百度文心一言 | ERNIE 4.5 Turbo, ERNIE 4.0 Turbo | ERNIE 4.5 Turbo: $0.003/M tokens, ERNIE 4.0 Turbo: $0.012/M | ERNIE 4.5 Turbo: $0.003/M, ERNIE 4.0 Turbo: $0.012/M | ✅ 国内直连 | ERNIE Speed/Lite/Tiny 免费调用,每个模型每月 10 万 tokens | 查看详情 → |
| #8 | Vercel AI Gateway | OpenAI GPT-4o / GPT-4o-mini / GPT-5.x, Anthropic Claude Opus 4.8 / Sonnet 4 | 免费层: $5/月 AI Gateway Credits;付费层: 按量付费,Provider 官价零加价 | Token 零加价(含 BYOK);附加功能另计 | ⚠️ 需代理(境内需 VPN) | 每团队每月 $5 免费 AI Gateway Credits(仅限免费层可用模型) | 查看详情 → |
| #9 | 月之暗面 Kimi | Kimi K3, Kimi K2.7 Code | K3: ¥2/M cached, ¥20/M uncached; K2.7/K2.6: ¥1.10-¥1.30/M cached, ¥6.50/M uncached | K3: ¥100/M; K2.7/K2.6: ¥27/M | ✅ 中国大陆直连 | 新用户 ¥15 代金券不可用于 Kimi K3;需充值后体验 | 查看详情 → |
| #9 | Mistral AI | Mistral Large 2, Mistral Small | Large 2: $2/M, Codestral: $1/M, Ministral 8B: $0.10/M | Large 2: $6/M, Codestral: $3/M, Ministral 8B: $0.10/M | ❌ 需代理 | 注册送免费 API 额度(La Plateforme 每日速率限制内免费) | 查看详情 → |
| #9 | Dify | GPT-4o / GPT-4o-mini / o1 / o3 (via OpenAI), Claude 3.5 Sonnet / Haiku / Opus (via Anthropic) | Sandbox 免费:200 GPT-4 调用 + 5MB vector 存储 + 10 文档 + 5,000 API 调用/月;Professional $59/workplace/月:100MB向量 + 100文档 + 50 工作流 | Team $159/workplace/月:500MB向量 + 500文档 + 200 工作流 + 500K 触发事件;Enterprise 合同制(SSO/私有部署/SLA) | ✅ 国内母公司 LangGenius,原生中文支持、阿里云一键部署、SOC 2/GDPR/ISO 27001 认证,国内云完全可用 | Sandbox 免费层:200 GPT-4 调用 + 5MB vector 存储 + 10 文档 + 2 工作流 + 5,000 API 调用/月(30 天日志保留) | 查看详情 → |
| #10 | 智谱 AI GLM | GLM-5.3, GLM-5.3-Flash | GLM-5.3 (Workers AI): $1.40/M + $0.26/M cached; GLM-5.2: ¥8/M; GLM-5.3-Flash (Workers AI): $0.15/M; GLM-5: ¥4-6/M; GLM-4.7: ¥2-4/M; GLM-4.5-Air: ¥0.8/M; GLM-4.7-Flash: ¥0/M (free) | GLM-5.3 (Workers AI): $4.40/M; GLM-5.2: ¥28/M; GLM-5.3-Flash (Workers AI): $0.50/M; GLM-5: ¥18-22/M; GLM-4.7: ¥8-16/M; GLM-4.5-Air: ¥2-8/M; GLM-4.7-Flash: ¥0/M (free) | ✅ 国内直连 | GLM-4-Flash 完全免费;注册送 500万 tokens 体验包 | 查看详情 → |
| #10 | Cohere | Command R+, Command R7 | Command R7: $2.50/M, Command R: $0.50/M, Embed: $0.10/M | Command R7: $10/M, Command R: $1.50/M | ❌ 需代理 | 免费 Trial API Key(2500次调用/月,限 Command R) | 查看详情 → |
| #10 | Cloudflare AI Gateway | 通过 Gateway 路由到 20+ Provider 任意模型 | 免费层: 100万次请求/月;超出 $0.20/100万次 + Provider 自身费用 | 按请求计费,不按 token 计;Gateway 本身加价 $0.20/100万请求 | ✅ 全球 CDN(国内边缘节点通过合作伙伴运营) | 每月 100 万次请求免费(Gateway 服务费) | 查看详情 → |
| #11 | 腾讯混元 | Hunyuan Hy4 preview, Hunyuan Turbo | Turbo S: ¥0.8/M tokens, Turbo: ¥1.2/M, Lite: ¥0.3/M | Turbo S: ¥0.8/M tokens, Turbo: ¥4.8/M, Lite: ¥0.3/M | ✅ 国内直连 | 新用户 100万 tokens 免费额度(180天有效) | 查看详情 → |
| #11 | Fireworks AI | Llama 3.3 70B, Firefunction-v2 | Firefunction-v2: $0.90/M, Llama 3.3 70B: $0.90/M, DeepSeek-R1: $2.00/M | Firefunction-v2: $0.90/M, Llama 3.3 70B: $0.90/M, DeepSeek-R1: $2.00/M | ❌ 需代理 | 新用户 $0.50 免费额度 | 查看详情 → |
| #11 | Portkey | OpenAI GPT-4o / GPT-5 / GPT-5.6 family, Anthropic Claude Opus 5 / Sonnet 5 / Haiku | Free 层: 10,000 请求/月;Hobby $49/月(10万请求);Growth $249/月(100万请求);Enterprise 合同制 | 请求量计费,不收 token 路由费;BYOK 免费;日志/可观测性/Caching 按订阅档位开放 | ❌ 需代理(portkey.ai 域名在国内访问不稳定,推荐自托管开源版本) | 每月 10,000 次请求;100K 日志保留;社区 Slack 支持;BYOK 免费;无信用卡要求 | 查看详情 → |
| #11 | Black Forest Labs | FLUX.2 [pro], FLUX.1.1 [pro] | FLUX.2 [pro]: $0.05/MP, FLUX.1.1 [pro]: $0.04/MP, FLUX.1 [schnell]: $0.003/MP, FLUX.1 Fill: $0.05/MP | Flat per-megapixel billing, no tier-based markup | ❌ 需代理(BFL API 部署在 AWS US-East + EU-Frankfurt,国内访问需稳定代理) | Free tier via API: limited generations per minute for [schnell] and [dev] endpoints; [pro] requires paid credits | 查看详情 → |
| #12 | 字节豆包 Doubao | Doubao-Seed-2.0 (旗舰推理), Doubao-Seed-2.0-lite (高性价比) | Seed-2.0: ¥0.8/M, Seed-2.0-lite: ¥0.5/M, Seed-2.0-mini: ¥0.15/M | Seed-2.0: ¥2/M, Seed-2.0-lite: ¥1/M, Seed-2.0-mini: ¥0.6/M | ✅ 国内直连 | 新用户 50万 tokens 体验包 | 查看详情 → |
| #12 | LiteLLM | 通过 Proxy 转发 100+ Provider 任意模型 | 开源免费自部署;LiteLLM Cloud 按量计费 | 开源免费自部署;LiteLLM Cloud 按量计费 | ✅ 开源自部署,无限制 | 开源版完全免费(自托管) | 查看详情 → |
| #12 | 硅基流动 | Qwen/Qwen3.5-Plus, Qwen/Qwen2.5-72B-Instruct | ¥0.4-2/1M tokens (Qwen2.5-7B ~¥0.4,Llama 3.3 70B ~¥2) | ¥0.4-2/1M tokens | ✅ 国内直连 | 注册送 ¥1-200 免费额度(按活动调整),足够跑 1-20M tokens | 查看详情 → |
| #12 | Jina AI | jina-embeddings-v3, jina-embeddings-v2-base-en | Embeddings v3: $0.02/M tokens (10K free); Reranker v3: $0.018/M tokens (10K free); Reader: $0.02/M tokens (free up to 1M tokens); CLIP-v1: $0.02/M tokens | Flat per-million-token billing, no tier-based markup; batch discount of 50% available on all endpoints | ❌ 需代理(Jina API 部署在 AWS US-East + EU-Frankfurt,国内访问需稳定代理) | 10,000 free tokens/month on Embeddings + Reranker; Reader API free up to 1M tokens/month; no credit card required to start | 查看详情 → |
| #12 | Arize Phoenix | OpenTelemetry 原生 trace collector,支持任意 LLM 框架, Phoenix Cloud 托管 + 自托管 (Apache-2 / Elastic-2.0) | Cloud 免费层: 10 GiB 存储/工作空间;自托管: $0(仅基础设施费) | Phoenix Cloud 付费: 按量付费(存储计费);Arize AX 企业: 合同制 | ✅ 开源自部署 + Cloud 托管(境外服务,境内需代理访问 SaaS) | Phoenix Cloud 免费套餐:每工作空间 10 GiB 存储(无时间限制) | 查看详情 → |
| #13 | Block AI(b.ai) | GPT-4o, GPT-4o-mini | 主流模型 + 0-10% 加价(如 GPT-4o $2.5→$2.75/M) | 主流模型 + 0-10% 加价 | ✅ 国内直连 | 注册送体验额度(每天免费请求次数待确认) | 查看详情 → |
| #13 | Groq | Llama 3.3 70B, Llama 3.1 405B | Llama 3.3 70B: $0.59/M, DeepSeek-R1: $0.89/M, Mixtral 8x7B: $0.15/M | Llama 3.3 70B: $0.79/M, DeepSeek-R1: $0.89/M, Mixtral 8x7B: $0.15/M | ❌ 需代理 | 免费速率限制(足够个人开发和小型测试使用) | 查看详情 → |
| #13 | fal.ai | FLUX.2 [pro] / [dev] / [schnell], Kling Video 2.1 / 2.0 | FLUX.2 [pro]: $0.05/MP (image), $0.08/sec (video); Kling 2.1: $0.10/sec; HunyuanVideo 1.5: $0.08/sec | Flat per-second / per-megapixel billing, no tier-based markup | ❌ 需代理(fal.ai 部署在 AWS US-East,国内访问需稳定代理) | $1 free credit on signup (no credit card required); 100+ models available free for first-time evaluation | 查看详情 → |
| #13 | Helicone | 通过 AI Gateway 代理 100+ Provider 任意模型, AI Gateway + LLM Observability 集成 | 免费层: 10,000 请求/月 + 1GB 存储;Pro $79/月(团队版含告警、HQL 查询);Team $799/月(SOC-2/HIPAA) | 按月订阅 + 用量计费(超过免费额度) | ✅ 开源自部署 + SaaS | 每月 10,000 次免费请求 + 1GB 存储 (Hobby) | 查看详情 → |
| #14 | FreeModel | DeepSeek V3 / R1, Qwen 2.5 / QwQ-32B | 模型相关(平台加价待确认) | 模型相关 | ✅ 国内可用(直连) | 注册送额度(具体金额待确认) | 查看详情 → |
| #14 | Cerebras | Cerebras Llama 3.3 70B, Cerebras Llama 3.1 405B | $0.10/M to $0.60/M tokens(所有模型统一 $0.60/M input) | $0.10/M to $0.60/M tokens(统一 $0.60/M output) | ❌ 需代理 | 免费速率限制(限制 RPM,适合测试) | 查看详情 → |
| #14 | AI21 Labs (Jamba) | Jamba 1.5 Mini, Jamba 1.5 Large | Jamba 1.5 Mini: $0.20/M tokens, Jamba 1.5 Large: $2/M tokens | Jamba 1.5 Mini: $0.20/M tokens, Jamba 1.5 Large: $2/M tokens | ❌ 需代理 | 无免费额度 | 查看详情 → |
| #14 | CoreWeave | NVIDIA HGX H100 ($49.24/hr), HGX H200 ($50.44/hr), HGX B200 ($68.80/hr), A100 80GB ($21.60/hr), L40S ($18.00/hr), L40 ($10.00/hr), GB200 NVL72 ($42.00/hr) — On-Demand/Spot GPU instances, Serverless Inference (W&B Inference, pay-per-token): GLM 5.2, Kimi K2.6/K2.7, DeepSeek R1/V3, Llama 3.x, Qwen — OpenAI-compatible API | GPU On-Demand hourly: NVIDIA HGX H100 $49.24, HGX H200 $50.44, HGX B200 $68.80, A100 80GB $21.60, L40S $18.00, L40 $10.00, GB200 NVL72 $42.00. Spot discounts: H100 $19.71, H200 $20.93, B200 $34.11, A100 $9.65. Serverless Inference is pay-per-token via W&B Inference; Dedicated Inference is GPU-hour / node-based. | ❌ 无中国大陆直连节点。数据中心集中在美国;2026 年已扩展至印尼(首个 APAC 节点)、英国(两个数据中心运营中)、瑞典。国内访问需代理,跨太平洋延迟较高。 | 无永久免费层。Serverless Inference 按 token 计费;GPU 实例按小时计费(On-Demand)或有 spot 折扣。0EM(零出口费迁移)计划在迁移阶段免除出口费。 | 查看详情 → | |
| #15 | APIKEY.FUN | GPT-4o, Claude 3.5 Sonnet | 主流模型,价格约为官方 API 的 60-80% | 主流模型,价格约为官方 API 的 60-80% | ✅ 国内直连 | 注册送少量免费额度 | 查看详情 → |
| #15 | DeepInfra | Meta-Llama-3.3-70B-Instruct, Meta-Llama-3.1-405B-Instruct | Llama 3.3 70B: $0.49/M, DeepSeek-V3: $0.99/M, DeepSeek-R1: $1.09/M | Llama 3.3 70B: $0.73/M, DeepSeek-V3: $0.99/M, DeepSeek-R1: $1.09/M | ❌ 需代理 | 注册即送免费额度(每日有限免费调用) | 查看详情 → |
| #15 | NVIDIA NIM | Nemotron-3 Ultra 550B (a55B), Nemotron-3 Super 120B (a12B) | Partner-dependent (Nemotron-3 Ultra 550B: $0.50-0.90/1M tokens) | Partner-dependent (Nemotron-3 Ultra 550B: $1.70-3.60/1M tokens) | ❌ 需代理(build.nvidia.com 平台为美国服务,受出口管制限制) | 免费开发端点(无需信用卡,生成 API Key 即可开始) | 查看详情 → |
| #15 | Amazon Bedrock | Claude Opus 4.8 / 4.7 / 4.6 / 4.5, Claude Sonnet 4.6 / 4.5 / 4 | Claude Sonnet 4.5: $6/M, DeepSeek V3.2: $0.62/M, Mistral Large 3: $0.50/M, GPT-5.5: $5.50/M, GPT-5.4: $2.75/M, Grok 4.3: $1.25/M, Qwen3 235B: $0.23/M, Kimi K2.5: $0.60/M | Claude Sonnet 4.5: $30/M, DeepSeek V3.2: $1.85/M, Mistral Large 3: $1.50/M, GPT-5.5: $33/M, GPT-5.4: $16.50/M, Grok 4.3: $2.50/M, Qwen3 235B: $0.91/M, Kimi K2.5: $3/M | ❌ 大陆用户需代理(AWS 中国区(由光环新网/西云数据运营)尚未提供 Bedrock 服务,需使用全球区域 + VPN/代理) | AWS Free Tier 赠送新用户最高 $200 抵扣金(6个月有效)。Bedrock 按用量抵扣,无永久免费额度 | 查看详情 → |
| #15 | Hyperbolic | DeepSeek R1, DeepSeek R1-0528 | Per-token inference scaled from market GPU rates; H100 from $2.89/GPU-hour, H200 $3.49/GPU-hour | Pay-as-you-go GPU compute billed per hour; reserved discounts for committed capacity | ❌ 需代理 | 无长期合同;按量计费按小时结算;无永久免费档,但无配额上限、即开即用 | 查看详情 → |
| #15 | Fireworks AI | DeepSeek V4 Pro — $1.74 input / $0.145 cached / $3.48 output per 1M tokens (Serverless Standard; Priority $2.61 / $0.218 / $5.22), DeepSeek V4 Flash (0731) — $0.22 input / $0.007 cached / $0.66 output per 1M (Standard) | Serverless (Standard, $/1M tokens): DeepSeek V4 Pro $1.74、V4 Flash $0.22、Kimi K3 $3.00、K2.7 Code $0.95、MiniMax M3 $0.30、Qwen 3.7 Plus $0.40、GPT OSS 120B $0.15、GLM 5.1 $1.40。嵌入按模型档 $0.008-$0.016/1M。按需 GPU 按小时:H100 $7.00、H200 $7.00、B200 $10.00、B300 $12.00、GB300 $18.00(9 月 1 日上调)。 | Serverless 输出($/1M):DeepSeek V4 Pro $3.48、V4 Flash $0.66、Kimi K3 $15.00、MiniMax M3 $1.20、Qwen 3.7 Plus $1.60、GPT OSS 120B $0.60、GLM 5.1 $4.40。缓存输入显著更低(如 DeepSeek V4 Pro $0.145)。训练:SFT 每 1M 训练 token $0.50(≤16B)到 $10.00(>300B);Serverless Training API 仅按 token 计费、无闲置 GPU 成本。 | ❌ 无中国大陆直连端点。作为美国平台,大陆直连需代理或海外中转,跨太平洋首字节延迟通常 150-250ms。最近 APAC 端点为东京多区域。 | 注册即享 $1 免费 serverless 额度(pay-per-token,postpaid billing);无永久免费档。按需 GPU 与训练按用量计费。 | 查看详情 → |
| #16 | Replicate | Llama 3.3 70B, DeepSeek-R1 | 模型调用按运行时长 + GPU 型号计费,$0.00025/s/A100起步 | 按运行时长计费(非 token 计费模式) | ❌ 需代理 | 新用户 $0.50 免费额度 | 查看详情 → |
| #16 | SambaNova | Llama 3.3 70B, Llama 3.1 405B | Llama 3.3 70B: $0.50/M, DeepSeek-R1: $0.80/M, Qwen 72B: $0.40/M | Llama 3.3 70B: $0.50/M, DeepSeek-R1: $0.80/M, Qwen 72B: $0.40/M | ❌ 需代理 | 免费速率限制(测试用途) | 查看详情 → |
| #16 | 零一万物 Yi | Yi-Lightning (智能路由 → DeepSeek-V3/Qwen3-30B-A3B/Yi-Lightning), Yi-Vision-v2 (视觉理解,路由 → Qwen2.5-VL-72B/Yi-Vision-V2) | ¥0.99/1M total tokens(Yi-Lightning,input+output 合并计费) | ¥0.99/1M total tokens(合并计费,不区分 input/output) | ✅ 国内直连(国内公司,支持手机号注册 + 支付宝/微信支付) | 注册即开通免费速率限制(有限 RPM/TPM),无初始赠送余额 | 查看详情 → |
| #16 | Nebius | DeepSeek V4 Flash, DeepSeek V4 Pro | Per-token OpenAI-compatible inference: DeepSeek-V4-Flash $0.14/M input, $0.28/M output; Kimi K3 $3/$15; MiniMax M3 $0.30/$1.20; Llama 3.3 70B $0.13/$0.40 | Two flavors per model - Base and Fast (-fast suffix) - same outputs, Fast trades higher token price for lower latency via speculative decoding | ❌ 需代理 | 注册即可起步,无月费;按量付费(pay-as-you-go)按 token 计费;弹性动态限流,随用量自动扩容至基础额度 20 倍,无需预购容量 | 查看详情 → |
| #17 | Anyscale | Llama 3.3 70B, Llama 3.1 405B | Llama 3.3 70B: $0.39/M, DeepSeek-R1: $1/M, Mixtral 8x22B: $0.90/M | Llama 3.3 70B: $0.39/M, DeepSeek-R1: $1/M, Mixtral 8x22B: $0.90/M | ❌ 需代理 | 新用户 $100 免费额度(一次性),另提供少量项目启动额度($3-$5) | 查看详情 → |
| #17 | DigitalOcean Gradient | Llama 3.3 70B, Llama 3.1 405B | Llama 3.3 70B: $0.50/M, DeepSeek-R1: $0.80/M | Llama 3.3 70B: $0.50/M, DeepSeek-R1: $0.80/M | ❌ 需代理 | 新用户 $100 免费额度(60 天有效) | 查看详情 → |
| #17 | MiniMax(海螺 AI) | MiniMax-01 (Lightning / Turbo / Pro), MiniMax-Text-01 | MiniMax-01 Lightning: ¥0.2/M tokens, Turbo: ¥0.8/M, Pro: ¥2/M, M3: ¥1.5/M | MiniMax-01 Lightning: ¥0.2/M tokens, Turbo: ¥0.8/M, Pro: ¥2/M, M3: ¥1.5/M | ✅ 国内直连 | 注册即送 100 万 tokens 免费额度(90 天有效期) | 查看详情 → |
| #18 | Stability AI | Stable Diffusion 3.5, Stable Diffusion XL | SD3.5: $0.0065/image, SDXL: $0.01/image, Stable Code: $0.50/M tokens | 按图像/视频/音频输出计费,非 token 计费 | ❌ 需代理 | 新用户 25 次免费生成 | 查看详情 → |
| #18 | Novita AI | Llama 3.3 70B, Llama 3.1 405B | Llama 3.3 70B: $0.59/M, DeepSeek-V3: $1.00/M | Llama 3.3 70B: $0.79/M, DeepSeek-V3: $1.00/M | ⚠️ 部分可用(新加坡节点,国内延迟尚可) | $0.50 免费试用额度 | 查看详情 → |
| #18 | Baseten | Llama 3.3 70B, Llama 3.1 405B | Per-second GPU billing (H100: $1.89/hr, A100: $0.83/hr, L40S: $0.54/hr) | Model-dependent (inference time × GPU rate) | ⚠️ 部分可用(需稳定代理连接) | $30 in inference credits on signup (valid 30 days); custom model hosting free tier: 1 deployed model | 查看详情 → |
| #18 | Mem0 | OpenAI GPT-4o / GPT-4-Turbo / GPT-3.5, Anthropic Claude 3.5 Sonnet / Haiku / Opus | Hobby 免费:10,000 memories + 1,000 retrieval API calls/月;Starter $19/月:50,000 memories + 5,000 calls | Pro $249/月:Unlimited memories + 50,000 retrieval API calls + Graph Memory + 多项目;Enterprise 合同制 | ✅ 开源(Apache-2.0,61k+ stars)+ Cloud SaaS;国内自部署无障碍 | Hobby 免费层:10,000 memories + 1,000 retrieval API calls/月(无限 end users),社区支持 | 查看详情 → |
| #18 | Tavily | Search (real-time web + AI summary, 1 credit/call), Extract (clean markdown from URL, 1 credit/page) | Free 1,000 credits/month(滚动 30 天窗口,无信用卡);Pay-as-you-go $0.008/credit(基础 Search 1 credit);Project 起步 $30/month 含 4,000 credits,$0.008/credit 超额;Growth $200/month 含 40,000 credits,$0.006/credit 超额;Pro $1,000/month 含 250,000 credits,$0.005/credit 超额;Enterprise 合同制(custom 配额 + SLA + 无日志合约) | 按 endpoint 计费:Search 1 credit/次;Extract 1 credit/页;Crawl 1 credit/页;Map 1 credit/页;Research 5 credits/次(替代 5-10 Search + 2-3 Extract + 1 合成 LLM);AI Extraction 1 credit/次 | ⚠️ Tavily API 部署在 AWS US-East / EU-West;国内直连延迟 200-400ms;推荐用 Cloudflare Worker 或腾讯云 Edge Function 中转(降至 50-100ms);无 CN 节点计划;无国内信用卡支付渠道 | Free 永久:1,000 credits/month(滚动 30 天窗口),无信用卡、无功能限制;Search/Extract/Crawl/Map/AI Extraction 各占独立配额(1,000 Search + 500 Crawl);超额返 HTTP 429,无超额计费 | 查看详情 → |
| #19 | Lepton AI | Llama 3.3 70B, Llama 3.1 405B | Llama 3.3 70B: $0.80/M; Llama 3.1 405B: $3.50/M; Qwen 2.5 72B: $0.80/M; DeepSeek-R1: $2.00/M | Same rate as input for most models (symmetric pricing) | ⚠️ 部分可用(ap-northeast-1 东京 region 距离最近,国内访问需代理) | $5 试用额度(30 天有效,足够完整评估任意一款主机模型) | 查看详情 → |
| #19 | 可灵 AI(快手) | Kling 3.0 (native 4K video, multi-shot sequencing, Native Audio), Kling 3.0 Omni (multimodal video, video/image reference input) | Per-second video billing. Kling 3.0 Turbo: 720p $0.112/s, 1080p $0.14/s. Kling 3.0: 720p $0.084-0.126/s, 1080p $0.112-0.168/s (Native Audio), 4K $0.42/s. Kling 3.0 Omni: 720p $0.084-0.126/s, 1080p $0.112-0.168/s, 4K $0.42/s, varies by video/audio input. | Image API: Kling Image 3.0 $0.028/image (1K/2K), 3.0-omni $0.028-0.056/image (up to 4K), Image 2.1 $0.014-0.028/image, Image O1 $0.028/image, Multi-Shot $0.07/call. Video list price 1 Unit = $0.14; image list price 1 Unit = $0.0035. | ✅ 官方支持部分中国区域调用(快手自研,国内云节点可用;海外 klingai.com 需网络环境) | 新用户注册后可获得少量免费体验积分(平台内赠送 credits),可体验部分视频/图片生成;正式 API 使用无永久免费套餐,按生成的秒数/张数计费。 | 查看详情 → |
| #20 | Hugging Face | Llama 3.3 70B, Llama 3.1 405B | 推理 API: $0.04-0.60/M tokens 取决于模型 | 推理 API: $0.04-0.60/M tokens | ❌ 需代理(huggingface.co 在国内被墙) | 免费推理 API(有限速率限制) | 查看详情 → |
| #20 | ElevenLabs | Eleven Multilingual v2, Eleven Turbo v2.5 | TTS: $0.03-0.30/1K chars depending on model/quality; STT: $0.47/hour; Sound Effects: $0.04/request | Same as input rate per model | ⚠️ 部分可用(需稳定代理) | 10,000 credits/month (≈10 minutes audio, free plan with limited features) | 查看详情 → |
| #20 | Ideogram | Ideogram 3.0, Ideogram Turbo | Ideogram 3.0: $0.04/image (standard), $0.08/image (Turbo) | 4 images per generation (default), additional images via paid credits | ❌ 需代理 | 新用户赠送 $1 体验金,约 25 次标准生成 | 查看详情 → |
| #20 | Lambda GPU Cloud | NVIDIA HGX B200 180GB (on-demand GPU instances; $6.69/GPU/hr), NVIDIA H100 SXM 80GB ($3.99/GPU/hr on-demand) | GPU instance on-demand per GPU/hr: NVIDIA B200 SXM6 $6.69, H100 SXM $3.99, A100 SXM 80GB $2.79, A100 SXM 40GB $1.99, GH200 $2.29, A100 PCIe $1.99, A6000 $1.09, A10 $1.29, V100 $0.79. 1-Click Cluster reserved per GPU/hr: B200 $9.86 (16 GPU) to $8.87 (256+), H100 $6.16 to $5.54. Superclusters / Private Cloud (4000+ GPUs) via sales. No egress fees. | Billed by the second per GPU instance (no idle GPU cost); clusters and reserved capacity carry a minimum commitment (2 weeks to 1 year). Managed orchestration: Managed Kubernetes and Slurm included at no extra per-node fee; Lambda Stack one-line install. No per-token inference pricing — Lambda's Inference API is winding down; self-hosted inference billed as GPU runtime. | ❌ 无中国大陆直连节点;数据中心在美国(加州等区域),国内访问需代理或海外中转。 | 无永久免费 GPU 层;注册后按用量计费(GPU 实例按秒计费)。1-Click Clusters 与 Superclusters 需预付 / 承诺周期(2 周至 1 年)。 | 查看详情 → |
| #21 | Deepgram | Nova-3 (STT, Monolingual), Nova-3 (STT, Multilingual) | STT: $0.0043-0.012/min; TTS: $0.015-0.030/1K chars | N/A | ⚠️ 部分可用(需代理,无中国大陆数据中心) | New users get $200 in free credits; no credit card required | 查看详情 → |
| #21 | RunPod | B300 288GB HBM3e, B200 180GB | Per-second GPU: H100 SXM $2.99/hr, A100 SXM $1.49/hr, B200 $5.89/hr, L40S $0.99/hr, RTX 4090 $0.69/hr | Same as input (per-GPU-second — no idle fees, no egress fees) | ⚠️ 部分可用(基础设施在全球 31 区域,国内直连不稳定,需稳定代理) | $5 starter credit on signup, no credit card required; per-second billing means spiky workloads never pay for idle | 查看详情 → |
| #21 | Luma AI | Ray 3.2 (flagship video, 1080p, multi-keyframe up to 16 frames, V2V up to 20s), Ray 3.14 (previous-gen video model) | Credit-based subscription. Plans: Plus $30/mo (10,000 credits), Pro $90/mo (40,000 credits), Ultra $300/mo (150,000 credits). Ray 3.2 video: 1080p text-to-video 400 credits/5s, 720p 100 credits/5s, Draft 20 credits/5s; Seedance 2.0 1080p 240 credits/sec, 4K 959 credits/sec. | Image credits: Uni-1 30 credits/image, Seedream 1-3 credits/image, GPT Image 2 from 3 (Low-1K) to 255 (High-4K) credits. Audio: ElevenLabs v3 TTS 21 credits/1,000 chars. Utilities: background removal 1 credit/image, reframe video 32 credits/sec, upscale to 4K 17 credits/sec. | ❌ 需代理(无中国大陆直连节点,API 托管于全球 AWS/GCP 边缘) | 无免费 API 套餐。Luma 提供有限免费信用额度供新用户体验(平台注册后少量免费 credits),但正式 API 使用需至少订阅 Plus $30/月(10,000 credits)。 | 查看详情 → |
| #22 | Chroma | All-MiniLM-L6-v2 (default, 384-dim, SBERT), All-mpnet-base-v2 (768-dim, higher quality SBERT) | Chroma OSS 完全免费(Apache 2.0),本地/嵌入式运行;Chroma Cloud Free 永久层:50K 向量、5K 查询/月、0.5 GB 存储;Pro $0.30/M 向量-月 + $0.01/M 查询(月最低 $10);Enterprise 合同制(自定义节点、HIPAA、SOC 2) | 存储(GB-月)+ 向量数 + 查询数 三维度计费;无 per-embedding API 费;Default Embedding Function 本地推理零成本;BYO Embedding 由第三方 API 计费 | ✅ Chroma OSS 可完全本地运行(国内服务器、零延迟);Chroma Cloud 节点在 AWS us-east-1 / eu-west-1 / ap-southeast-2(2026 H2 新加坡节点待上线),直连大陆延迟 200-400ms | Chroma OSS 永久免费、Apache 2.0、无功能限制、无向量数限制(本地硬件决定);Chroma Cloud Free 永久层:50K 向量 + 5K 查询/月 + 0.5 GB 存储 + 1 个 Collection + 社区支持 | 查看详情 → |
| #22 | Voyage AI | voyage-4-large, voyage-4 | voyage-4-large: $0.12/M, voyage-4: $0.06/M, voyage-4-lite: $0.02/M, voyage-context-3: $0.18/M, voyage-code-3: $0.18/M, voyage-multimodal-3.5: $0.12/M | (embedding API — no output tokens; rerank-2.5: $0.05/M tokens, rerank-2.5-lite: $0.02/M tokens) | ⚠️ 部分可用(官网在中国可访问,但 API 在中国大陆的稳定性未公开声明) | 200M free tokens per account (text embedding + reranker), 50M free tokens for legacy domain models, 200M free text tokens + 150B free pixels for multimodal | 查看详情 → |
| #22 | Runway | Gen-4 Turbo, Gen-4 Alpha | Gen-4 Turbo: 50 credits/5s clip (~$0.50), Gen-4 Alpha: 100 credits/5s (~$1.00), Act-Two: 150 credits/5s (~$1.50), Frames: 200 credits/sequence (~$2.00) | Credit-based system (~$0.01/credit). Credit packs: $12 (1,000 credits) to $200 (25,000 credits). Enterprise: custom pricing. | ❌ 需代理(AWS US-East + GCP US-Central 托管) | 无免费 API 套餐。最低 $12 购买 1,000 credits,约 20 次 Gen-4 Turbo 生成。Web 应用免费但无 API 访问。 | 查看详情 → |
| #22 | TwelveLabs | Pegasus 1.5 (flagship multimodal video understanding: frames + audio + speech + on-screen text), Pegasus 1.2 (previous-gen, lower-cost option: frames + audio + speech) | Free plan: 600 minutes video indexing + API usage (Pegasus 1.2 $0.042/min indexing one-time, $0.021/min input video, $0.0075/1k tokens output text, $0.0015/min monthly embedding infrastructure). Developer plan: paid overage, 90-day index retention lifted to unlimited. | Same per-minute and per-1k-tokens billing. Marengo embeddings billed separately by indexed minute. | ❌ 需代理(无中国大陆直连节点,API 托管于 AWS us-west-2 / us-east-1 区域) | 免费计划:注册即获 600 分钟视频索引额度(一次性累积额度),可调用全部核心 API(Pegasus 1.2/1.5、Marengo 3.0、Search、Embed、Analyze & Segment)。索引数据保留 90 天;升级 Developer 计划后索引保留变为无限。 | 查看详情 → |
| #22 | Mixedbread AI | Toast 1 — specialized search model (2026-08-13): matches/outperforms Claude Opus 5 and GPT-5.6 Sol on knowledge work, up to 10x cheaper and 12x faster; input $0.50 / cached $0.06 / output $1.20 per 1M LLM tokens (launch $0.30 / $0.036 / $0.72), mxbai-embed-large-v1 — 1,024-dimension open embedding model (Apache-2.0), 50M+ downloads across the mxbai family | 按量计费,三块:索引(Fast $1.50/百万内容 token;High Quality 含 OCR/转写/摘要/多模态增强 $3/百万)、搜索(Grep $0.10/千次、Semantic $4/千次、Toast 1 $1/千次,重排 +$3.50 或 +$1.50/千次)、存储 $0.50/百万内容 token/月。Toast 1 按 LLM token:输入 $0.50/百万(发布价 $0.30)。 | Toast 1 输出 $1.20/百万 LLM token(发布价 $0.72,40% 折扣);缓存输入 $0.06/百万(发布价 $0.036),缓存写入免费。Semantic 检索重排加价 $3.50/千次。 | ❌ 无中国大陆直连端点。Mixedbread 是美国/欧洲平台(法人实体 mixedbread ai inc.,柏林/旧金山),大陆直连需代理或海外中转,跨太平洋首字节延迟通常 150-250ms。 | Starter 免费送 $5 一次性额度,无需信用卡;3 个工作区用户、10 个 Store、每分钟 100 次请求,适合评估真实语料。无永久免费档。 | 查看详情 → |
| #23 | Crusoe Cloud | DeepSeek V3 0324 (legacy frontier: $0.50 input / $1.50 output per 1M tokens), DeepSeek V4 Flash (efficient frontier: $0.14 input / $0.28 output per 1M tokens, 70B tier) | Pay-as-you-go per 1M tokens, four tiers by parameter count: <16B ($0.40) / 16B-70B ($2.50) / 70B-300B ($6.00) / >300B ($10.00). Serverless Fine-Tuning follows identical 4-tier structure. Managed Inference spans DeepSeek V3/V4 Pro/V4 Flash, GLM 5.1/5.2, GPT-OSS 20B/120B, Gemma 4 31B-it, Kimi K2.6, Llama 3.1/3.3, Nemotron 3 family (VoiceChat/Ultra/Lightning/Nano/Super/Nano Omni), Qwen3 family (8B/235B/Qwen3.5/Qwen3.6), Yutori n1.5. Cached tokens at $0.03-$1.50 per 1M. | Self-Serve Deployments: NVIDIA H100 80GB HGX $5.50/hr, NVIDIA H200 141GB HGX $6.00/hr (dedicated endpoints for open + fine-tuned models). Tailored Deployments and Provisioned Throughput via sales engagement. Managed Kubernetes $0.10/cluster hour; Container Registry $0.10/GiB-month; Object Storage $0.06/GiB-month. | ❌ 无中国大陆直连节点。托管推理 API 托管于 Crusoe 美国数据中心(Colorado / Texas 等);国内访问需代理。 | 无永久免费层;按使用量计费;可联系销售获取批量折扣。Serverless Fine-Tuning 四档起价 $0.40/M tokens。 | 查看详情 → |
| #23 | Letta | Letta Agents SDK — open-source stateful agent runtime (Apache-2.0), with first-class persistent memory (core/archival/recall blocks) and runtime memory editing, Letta Code (`@letta-ai/letta-code`, Node.js 22.19+) — memory-first CLI coding agent, #1 on Terminal-Bench (Dec 2025 launch) | 按用量:Free $0(最多 3 个有状态 agent、BYOK)、Pro $20/月(最多 20 个 agent、Letta Auto 周/月额度 + 按量超额)、Teams Pro 按席位(共享 agent、权限)、Developer Plan 纯 API 调用按信用(无 agent 上限)、Enterprise 商务报价(自托管/定制模型)。服务端工具按 CPU 时间 $0.00015/秒计费。 | Letta Auto 用量超出 Pro 套餐限额后按 pay-as-you-go 模型价(参考 platform.letta.com/models)。Mods、Skills、Conversations 等功能均包含在套餐内,不单独计费。远程 MCP 工具由 MCP 供应商执行,不消耗 Letta 信用。 | ❌ 无中国大陆直连端点。Letta, Inc. 注册于美国旧金山,platform.letta.com 是单一全球端点;大陆生产流量需要代理或中转,跨太平洋首字节延迟通常 150-250ms。Agent harness 完全开源(Apache-2.0),可自托管在自有基础设施上规避直连限制。 | Free $0/月(最多 3 个有状态 agent),自带 API key;Letta Auto 用量与 Skills 受限,可用于评估和小型个人助手。 | 查看详情 → |
| #24 | Weaviate | Snowflake Arctic-embed-m-v1.5, Snowflake Arctic-embed-m-v2.0 | Free 永久层:100,000 对象 + 1 集合 + 7 天备份;Flex 月付 $45 起 + 从 $0.00465/1M 维度 + $0.12/GiB 存储;Premium 预付从 $400/月 | Dedicated 合同制 from $400/月 起;Vector dimension、Storage、Backup 三维度独立计费;Embeddings 按 token 用量计费 | ⚠️ 服务器在 AWS/GCP 海外节点,直连中国大陆延迟较高 | Free 永久层:100,000 对象、1 集合、7 天备份;2,000 Embedding 请求/日;Query Agent 1,000 请求/月 | 查看详情 → |
| #25 | Writer | Palmyra X5, Palmyra X4 | Palmyra X5: $0.60/M, Palmyra X4: $2.50/M | Palmyra X5: $6.00/M, Palmyra X4: $10.00/M | ⚠️ 部分可用(官网在中国可直接访问,但 API 稳定性未公开声明) | 14-day free trial, no credit card required (Starter plan); no standalone free API credits | 查看详情 → |
| #25 | Reka | Reka Core — top-tier multimodal model (text + image + audio + video); input $2.00 / output $6.00 per 1M tokens, image $0.02 / video $0.08 / audio $0.02 per minute, Reka Flash — cost-efficient model for everyday tasks (text + image + audio + video); input $0.80 / output $2.00 per 1M tokens, image $0.01 / video $0.06 / audio $0.015 per minute | Reka Edge $0.10/M 输入;Reka Flash $0.80/M;Reka Core $2.00/M;Reka Flash Research 按请求计费($25-$60/千次)。按量付费,无最低承诺。 | Reka Edge $0.10/M;Reka Flash $2.00/M;Reka Core $6.00/M。多模态(图像、视频、音频)单独按"每分钟"计费,Flash 模型视频 $0.06/分钟、Core $0.08/分钟;图像按调用计费(Flash $0.01、Core $0.02)。 | ❌ 无中国大陆直连端点。Reka 注册于美国加州 Sunnyvale(530 Lawrence Expressway),api.reka.ai 是单一全球端点;大陆生产流量需代理或中转,跨太平洋首字节延迟通常 150-250ms。开源权重可自托管(NVIDIA GPU ≥24GB VRAM),规避直连限制。 | 通过 platform.reka.ai 注册即可获取 API key;Pay-as-you-go 用量计费,不强制最低承诺,可在充值额度内自由试用。无永久免费档、无月赠送额度。 | 查看详情 → |
| #26 | Qdrant | FastEmbed: BAAI/bge-small-en-v1.5, FastEmbed: BAAI/bge-base-en-v1.5 | Free 永久层:1 GB 存储 + 0.5M 向量,无限期;Cloud Standard 月付 $25 起(预付 $250/yr)按存储 $0.06/GB-月;Cloud Pro 月付 $80 起 按存储 $0.04/GB-月 + 高 IOPS;Dedicated 合同制 $2,500/月起 | 按存储维度(GB)与可选性能包(Power Tiers)计费;无 API 调用费、无 per-vector 嵌入费;FastEmbed 按 token 数本地推理;BYO Embedding 由第三方 API 计费 | ⚠️ Cloud 运行在 AWS Frankfurt / N. Virginia / Sydney 等海外节点,直连中国大陆延迟较高;OSS 可在腾讯云/阿里云国内区域自托管 | Free 永久层:1 GB 存储 + 0.5M 1015-dim 向量;2 CPU/0.5 GiB RAM 实例;无限 API 请求;Community Discord 支持 | 查看详情 → |
| #27 | fal.ai | FLUX.2 Pro (text-to-image, $0.03/MP first MP + $0.015/extra MP), FLUX1.1 [pro] (ultra-fast text-to-image, per-megapixel) | 按请求/按时长计费,无月费。图像生成:FLUX.2 Pro $0.03/第一个 MP + $0.015/额外 MP、FLUX1.1 [pro] 按 MP、Nano Banana 2 $0.08/图、FLUX.1 [schnell] 无定价显示。视频:Kling v3 $0.084-0.168/秒、Seedance 2.0 $0.3034/秒(720p)。TTS:通过 ElevenLabs/MiniMax 等第三方提供。余额模式:预付 $10+ 存入信用余额,用完为止。开源模型按秒计费,秒级颗粒度。 | ✅ 大陆直连可用(fal.ai 使用全球 CDN,无需代理);延迟因模型推理节点而异,典型 200-500ms | 无免费 API 层(预付 $10 起买信用余额);无月费、无订阅、无隐藏费用;Serverless 模式下零运行时自动缩零 | 查看详情 → | |
| #28 | Aleph Alpha | Pharia-1-LLM-7B-control, Pharia-1-LLM-7B-control-aligned | Enterprise contract — no public per-token price (quote-based; PhariaAI on-prem deployment custom-priced) | Same — enterprise quote, no public per-token list | ❌ 不适用(德国主权 AI 厂商,受欧盟出口管制约束) | No public free tier — self-serve signup creates an account, but pricing is enterprise contract only | 查看详情 → |
| #28 | Cartesia | Sonic (real-time TTS), Sonic 2 (multilingual, voice cloning) | Sonic: $0.03/1K chars (Standard), $0.06/1K chars (Pro), $0.12/1K chars (Turbo); Sonic 2: $0.04-0.15/1K chars | TTS character-based pricing (output = audio), per-second for streaming | ⚠️ 部分可用(需稳定代理) | 10,000 characters/month free tier (≈10 minutes), no credit card required | 查看详情 → |
| #28 | Hume AI | EVI 3 (Empathic Voice Interface 3), OCTAVE TTS (voice design + cloning) | EVI 3: ~$0.096/min (pay-per-second); OCTAVE TTS: $0.048/1K chars; Expression Measurement: $0.0008/sec | Same as input (most endpoints symmetric) | Proxy required (US export controls) | Free credits on signup (amount per account), enough to trial EVI 3 | 查看详情 → |
| #28 | Modal | B300 / B200 / H200 / H100 / A100 / L40S / A10 / L4 / T4 (GPU), Self-hosted LLM inference (vLLM, SGLang, TensorRT-LLM) | 按 GPU 秒计费:H100 $0.001097/s、A100 80GB $0.000694/s、T4 $0.000164/s | 同输入(按 GPU 时间,无 idle 费用) | ❌ 需代理(基础设施在 AWS/GCP 海外区域,国内直连不稳定) | Starter 计划免费 + $30/月 计算额度(每月刷新) | 查看详情 → |
| #30 | AssemblyAI | Universal-3.5 Pro (Async, 18 languages, native code switching, best diarization), Universal-2 (Async, 99 languages, balanced accuracy) | 按音频小时计费:Universal-3.5 Pro $0.21/hr、Universal-2 $0.15/hr(pre-recorded) | N/A(speech-to-text 计费基于音频时长,无 token 概念) | ❌ 需代理(基础设施在 US,国内直连不稳定) | 免费账户每月 185 小时 pre-recorded + 333 小时 streaming,无需信用卡 | 查看详情 → |
| #30 | Pinecone | Serverless (Standard / Enterprise), Pod-based (s1 / p1 / p2) | Serverless Standard: $50/月最低,$4-$4.50/M读单元,$16-$18/M写单元,$0.0005/ingestion单元 | Serverless Enterprise: $500/月最低,$6-$6.75/M读,$24-$27/M写,$0.001/ingestion单元(multi-modal);Pod-based按节点小时计费 | ⚠️ 国内访问 pinecone.io 需代理(SaaS-only,无开源版) | 免费 Starter 计划(1个项目,Serverless索引,$50信用额度/月,有效期1年) | 查看详情 → |
| #39 | Together AI | DeepSeek V4 Pro, Qwen3.6-Plus / Qwen3.7-Max | $0.03-$1.74/MTok(模型相关) | $0.12-$4.50/MTok(模型相关) | ⚠️ 需代理(美国服务商) | 注册赠送 $1 初始额度(需绑定支付方式) | 查看详情 → |
| # | FriendliAI | GLM-5.2, GLM-5.1 | GLM-5.2/5.1: $1.4/M, DeepSeek-V3.2: $0.5/M, Qwen3-235B: $0.2/M, MiniMax-M2.5: $0.3/M, Gemma-4-31B: $0.14/M | GLM-5.2/5.1: $4.4/M, DeepSeek-V3.2: $1.5/M, Qwen3-235B: $0.8/M, MiniMax-M2.5: $1.2/M, Gemma-4-31B: $0.4/M | ⚠️ Partially available — servers in US & Korea, ~200-400ms latency from CN | No free tier — per-token billing with no minimum | 查看详情 → |
| #40 | Suno | v5.5 (latest model, royalty-free music), v5 (premium model) | Free: $0 (10 songs/day, non-commercial); Pro: $10/mo (500 songs, commercial rights); Premier: $30/mo (2,000 songs, Studio) | Annual: Pro $96/yr (20% off), Premier $288/yr (20% off); credit top-ups available as add-ons | ⚠️ 国内访问受限(需稳定海外代理),无官方中国大陆节点 | 免费版每天 10 首歌曲(不可商用),新账户直接开通,无需信用卡 | 查看详情 → |
| #40 | Liquid AI | LFM2.5-8B-A1B (MoE, 8B total / 1.5B active, 128K context) — open weights, free download / run / fine-tune under the LFM Open License (royalty-free until $10M annual revenue), LFM2.5-2.6B (dense, 128K context, agentic tool calling) — DSpark draft variant up to 3.18× GPU / 2.87× on-device decode speedup | 无按 token 云 API 定价——LFM 权重免费下载/运行/微调(LFM Open License,公司年收入超过 1000 万美元前免版税)。推理要么在自有硬件上自托管,要么通过 LEAP SDK 本地部署;无 per-token 请求费用。 | 无 per-token 输出计费。模型在 CPU/GPU/NPU 上本地或自托管运行,适合对延迟与隐私敏感的高频或离线工作负载(Liquid 的定位:「为何为云 API 按 token 付费,而非本地推理」)。 | ⚠️ 无托管云 API,因此无中国直连端点问题。模型权重自由下载并可在本地/自托管运行,包括在中国境内(对数据主权友好)。不需要云 API 网关或代理;token 处理完全在本地硬件上进行。 | 全部 14+ 个 LFM 模型在 LFM Open License 下免费下载/运行/微调——包括商用(公司年收入 ≤ 1000 万美元)。研究、教育与非营利使用永久免费、无收入上限。 | 查看详情 → |
| # | Nscale | Moonshot AI Kimi K2.5 — open chat model via serverless inference; per-1M input/output token pricing (verified on nscale.com/product/inference featured model list), Alibaba Cloud Qwen3 8B / Qwen3 4B Thinking / Qwen3 4B Instruct — open chat models with reasoning variants; per-1M token pricing | Nscale 采用充值余额(credit)按量计费,价格按模型分列,在 Nscale 控制台 AI Services → Models 与 /v1/models API 中按每 1M 输入 token 报价(文本模型);图像模型输入按 0 计费。定价结构经 serverless-openapi.yaml ModelPricing schema 验证(input = 每 1M 输入 token 美元价)。 | 输出按每 1M 输出 token 计费(文本模型);图像模型按每百万像素计费(ModelPricing.output 定义为 per-million output pixels)。具体美元单价在控制台与 /v1/models 端点内按模型展示,注册充值后可查。 | ❌ 无中国大陆直连端点。inference.api.nscale.com 是单一全球端点,数据中心集中于欧洲(英国 Loughton、挪威 Glomfjord/Narvik)与美国(德州、西弗吉尼亚);大陆生产流量需代理或中转,跨大西洋/跨太平洋首字节延迟通常 200-350ms。主打 EU 数据主权与在国境内处理,适合欧洲合规场景而非中国区低延迟。 | 无永久免费档。新用户注册后通过 Google SSO 登录,官方提示早期用户可能获得免费信用额度(promotional offer,非保证)。默认需先充值(credit)才能调用 serverless 推理 API。 | 查看详情 → |
| #41 | Morph | Kimi K3 (2.8T MoE) — flagship Fable-tier open-weight model at ~100 tok/s with 1M context; per-1M input/output token pricing (verified on docs.morphllm.com llms.txt Fast Models table 2026-08-28), GLM-5.2 (744B MoE) — Opus-tier open-weight model with 1M context; OpenAI `service_tier` parameter supported for default/standby processing | 按模型分别计费,按每 1M 输入 token 计(文本模型)。从 llms.txt 验证(2026-08-28):Kimi K3 $2.90、Qwen 3.5 397B $0.50、GLM-5.2 744B $1.10、GLM-5.3-Flash $0.15、MiniMax M3 $0.30、MiniMax M2.7 $0.279、DeepSeek V4 Flash (beta) $0.12、Qwen 3.8/3.6 27B $0.289、Gemma 4 31B $0.14。Qwen 3.5 397B 缓存输入 $0.30 per 1M(其他模型无独立缓存价)。Fast Apply 用 `auto` 模型按 prompt+completion 两档计(OpenRouter 验证 morph-v3-fast $0.0008/$0.0012 per 1k、morph-v3-large $0.0009/$0.0019 per 1k),Compact、Reflexes、WarpGrep 与 Router 按事件计费(详见 pricingEN)。 | 按模型分别计费,按每 1M 输出 token 计(文本模型)。从 llms.txt 验证:Kimi K3 $14.00、Qwen 3.5 397B $3.50、GLM-5.2 744B $4.10、GLM-5.3-Flash $0.50、MiniMax M3 $1.20、MiniMax M2.7 $1.20、DeepSeek V4 Flash (beta) $0.278、Qwen 3.8/3.6 27B $2.40、Gemma 4 31B $0.40。Fast Apply 输出按 morph-v3-fast $1.20 / morph-v3-large $1.90 per 1M token;Fast Apply 用 `auto` 推荐档(5,000-10,500 tok/s, ~98% 精度)按实际路由模型分档计费。Kimi K3 `service_tier: standby` 走 GLM-5.2 同等待遇,标准按 token 计。 | ❌ 无中国大陆直连端点。base URL https://api.morphllm.com 为单一全球端点,Y Combinator 背书的美国总部(AutoInfra, Inc.);OpenRouter 上 morph/morph-v3-fast 和 morph/morph-v3-large 走多区域路由,但官方端点不在大陆。中国大陆生产流量需代理或中转,跨太平洋首字节延迟通常 200-400 ms。 | 无永久免费档(无零元 starter plan)。新用户注册即可获得 API key 并按 pay-as-you-go 充值;官方为初创公司提供 Startup Credits(最高 $5,000),但需单独联系销售申请。Morph 主页明确「Start free. Get API Key」,但实际计费从首次 token 调用开始。 | 查看详情 → |
| #42 | Firecrawl | Scrape (v2 /scrape) — convert any URL into clean markdown / HTML / structured JSON via the JSON-mode schema option; 1 credit per page, Crawl (v2 /crawl) — recursively crawl a website and return content for every linked page; 1 credit per page scraped | 纯 credit-based 按 endpoint 与特性消耗,每月按 plan 重置额度;额外 credit 按 $5 一档购买(每档 1k–5k credits 不等)。从 firecrawl.dev/pricing 与 docs.firecrawl.dev/billing.md 验证(2026-08-29):Free 1k/月 $0;Hobby 5k/月 $19(年付 $16);Standard 100k/月 $83(年付 $99);Growth 500k/月 $333(年付 $399);Scale 1M/月 $599;Enterprise Custom。Credit 单价:Scrape / Crawl / Map 1 cr/page、Search 2 cr/10 results、Monitor 1 cr/page/check、Interact 2–7 cr/browser-minute。Batch Scrape / Extract 按子操作同价。Agent 5 daily runs free,超量动态计费。 | 无独立 output 维度——Firecrawl 输出的是 scraped markdown / HTML / JSON 内容,计费仅按输入端的 credit 消耗与并发浏览器占用。出站数据本身不另收费;唯一例外是 Search 在 ZDR(zero-data-retention)企业档下按企业价另议。 | ❌ 无中国大陆直连端点。Firecrawl 总部位于美国旧金山(YC W23),API base URL https://api.firecrawl.dev 单区域部署;docs.firecrawl.dev/billing.md 未列出任何亚太 / 中国路由层。中国大陆生产流量需代理或中转,跨太平洋首字节延迟通常 200-400 ms;Free 档 1000 credits 评估期可代理试用,但稳定流量建议自建代理或自部署 Firecrawl 自托管版(self-hosted Docker / Kubernetes)。 | Free 档永久免费:1,000 credits / 月,无需信用卡,支持 Scrape / Crawl / Map / Search / Batch Scrape / JSON mode 全部 endpoints;2 concurrent browsers;50,000 max queued jobs。适合评估与轻量使用,但 credit 重置周期每月初清零、当月用完不再补。Agent 5 daily runs free 适用于所有 plan(含 Free)。 | 查看详情 → |