找到最划算的 AI Token 方案

收录 OpenAI、Claude、DeepSeek、Gemini 等10+主流 AI 厂商的最新 API 价格,实时更新,免费对比,帮你省钱。

10+
收录厂商
50+
覆盖模型
实时更新
Updates
🇨🇳

国内可用 AI API

国内直连

DeepSeek(深度求索)

#4

DeepSeek

输入 DeepSeek-V3: ¥0.14/M token (≈$0.02), R1: ¥0.14/M
输出 DeepSeek-V3: ¥0.28/M token (≈$0.04), R1: ¥0.28/M
✅ 国内直连 免费额度

注册送 ¥500 免费额度(约500万tokens)

美团 LongCat

#6

Meituan LongCat

输入 OpenRouter: $0.60/M token (48B active MoE)
输出 OpenRouter: $2.40/M token (48B active MoE)
✅ Domestic self-host + OpenRouter (proxy) 免费额度

MIT licensed weights on Hugging Face; self-host on 8x H100+ GPU cluster

阶跃星辰(Stepfun)

#6

Stepfun

输入 Step-2: ¥6/M, Step-2-mini: ¥1/M, Step-1: ¥4/M, Step-R: ¥8/M
输出 Step-2: ¥18/M, Step-2-mini: ¥3/M, Step-1: ¥12/M, Step-R: ¥24/M
✅ 国内直连(platform.stepfun.com 端点,国内BGP优化) 免费额度

新用户 100万 tokens 免费额度(注册后30天有效)

阿里云百炼

#7

Alibaba Cloud Bailian

输入 Qwen3.5-Max: ¥4/M, Qwen3.5-Plus: ¥2/M, Qwen3.5-72B: ¥4/M, QwQ: ¥2/M
输出 Qwen3.5-Max: ¥12/M, Qwen3.5-Plus: ¥6/M, Qwen3.5-72B: ¥12/M, QwQ: ¥8/M
✅ 国内直连 免费额度

新用户 100万 tokens 免费额度(90天有效)

百度文心一言

#8

Baidu ERNIE (文心一言)

输入 ERNIE 4.5 Turbo: $0.003/M tokens, ERNIE 4.0 Turbo: $0.012/M
输出 ERNIE 4.5 Turbo: $0.003/M, ERNIE 4.0 Turbo: $0.012/M
✅ 国内直连 免费额度

ERNIE Speed/Lite/Tiny 免费调用,每个模型每月 10 万 tokens

月之暗面 Kimi

#9

Moonshot AI Kimi

输入 K3: ¥2/M cached, ¥20/M uncached; K2.7/K2.6: ¥1.10-¥1.30/M cached, ¥6.50/M uncached
输出 K3: ¥100/M; K2.7/K2.6: ¥27/M
✅ 中国大陆直连 免费额度

新用户 ¥15 代金券不可用于 Kimi K3;需充值后体验

智谱 AI GLM

#10

Zhipu AI GLM

输入 GLM-5.3 (Workers AI): $1.40/M + $0.26/M cached; GLM-5.2: ¥8/M; GLM-5.3-Flash (Workers AI): $0.15/M; GLM-5: ¥4-6/M; GLM-4.7: ¥2-4/M; GLM-4.5-Air: ¥0.8/M; GLM-4.7-Flash: ¥0/M (free)
输出 GLM-5.3 (Workers AI): $4.40/M; GLM-5.2: ¥28/M; GLM-5.3-Flash (Workers AI): $0.50/M; GLM-5: ¥18-22/M; GLM-4.7: ¥8-16/M; GLM-4.5-Air: ¥2-8/M; GLM-4.7-Flash: ¥0/M (free)
✅ 国内直连 免费额度

GLM-4-Flash 完全免费;注册送 500万 tokens 体验包

腾讯混元

#11

Tencent Hunyuan

输入 Turbo S: ¥0.8/M tokens, Turbo: ¥1.2/M, Lite: ¥0.3/M
输出 Turbo S: ¥0.8/M tokens, Turbo: ¥4.8/M, Lite: ¥0.3/M
✅ 国内直连 免费额度

新用户 100万 tokens 免费额度(180天有效)

字节豆包 Doubao

#12

ByteDance Doubao

输入 Seed-2.0: ¥0.8/M, Seed-2.0-lite: ¥0.5/M, Seed-2.0-mini: ¥0.15/M
输出 Seed-2.0: ¥2/M, Seed-2.0-lite: ¥1/M, Seed-2.0-mini: ¥0.6/M
✅ 国内直连 免费额度

新用户 50万 tokens 体验包

零一万物 Yi

#16

01.AI Yi (零一万物)

输入 ¥0.99/1M total tokens(Yi-Lightning,input+output 合并计费)
输出 ¥0.99/1M total tokens(合并计费,不区分 input/output)
✅ 国内直连(国内公司,支持手机号注册 + 支付宝/微信支付) 免费额度

注册即开通免费速率限制(有限 RPM/TPM),无初始赠送余额

MiniMax(海螺 AI)

#17

MiniMax (Hailuo AI)

输入 MiniMax-01 Lightning: ¥0.2/M tokens, Turbo: ¥0.8/M, Pro: ¥2/M, M3: ¥1.5/M
输出 MiniMax-01 Lightning: ¥0.2/M tokens, Turbo: ¥0.8/M, Pro: ¥2/M, M3: ¥1.5/M
✅ 国内直连 免费额度

注册即送 100 万 tokens 免费额度(90 天有效期)

🌍

国际 AI API

需代理

OpenAI

#1

OpenAI

输入 GPT-5: $1.25/M, GPT-5-mini: $0.25/M, GPT-4o: $2.50/M, GPT-4o-mini: $0.15/M
输出 GPT-5: $10/M, GPT-5-mini: $2/M, GPT-4o: $10/M, GPT-4o-mini: $0.60/M
❌ 需代理(受出口管制限制,国内用户需使用代理访问) 免费额度

新账户赠 $5 额度,有效期 3 个月

Anthropic Claude

#2

Anthropic Claude

输入 Sonnet 5 (intro): $2/M, Opus 5: $5/M, Sonnet 5/4.5: $3/M, Opus 4.8: $5/M, Haiku 4.5: $0.80/M
输出 Sonnet 5 (intro): $10/M, Opus 5: $25/M, Sonnet 5/4.5: $15/M, Opus 4.8: $25/M, Haiku 4.5: $4/M
❌ 需代理(受美国出口管制,国内用户需代理) 免费额度

Claude 免费版有限额度(网页端);API 端无赠送额度;Sonnet 5 限时价 $2/M(至 2026-08-31)

Azure OpenAI(微软 Azure)

#2

Azure OpenAI

输入 GPT-4o: $2.50/M, GPT-4o-mini: $0.15/M, o1: $15/M, o3: $10/M
输出 GPT-4o: $10/M, GPT-4o-mini: $0.60/M, o1: $60/M, o3: $40/M
⚠️ 世纪互联运营版(国内可用,模型比国际版少) 免费额度

新用户 $200 免费额度(30天有限)

xAI Grok

#4

xAI Grok

输入 Grok 4.6: $2/M tokens (fast variant $4/M)
输出 Grok 4.6: $6/M tokens (fast variant $12/M)
❌ 需代理 免费额度

无免费 API 额度(Grok Build 试用;初期免信用卡)

Gemini 多模态

#4

Gemini Multimodal

输入 Omni Flash: $0.15/M input tokens (multimodal); Nano Banana 2 Lite: $0.034/1K image tokens; Veo 3.1: $0.10/sec video
输出 Omni Flash: $0.60/M output tokens (multimodal); Nano Banana Pro: $0.12/1K image tokens
❌ 需代理(Google AI Studio 在中国大陆不可直接访问,需稳定代理) 免费额度

Free tier: Gemini Omni Flash 15 RPM, Veo 3.1 2 generations/day, Imagen 4 10 images/day — no credit card required

Google Gemini

#5

Google Gemini

输入 2.5 Pro: $1.25-10/M, 2.5 Flash: $0.15-1.25/M, 2.0 Flash: $0.10/M
输出 2.5 Pro: $5-40/M, 2.5 Flash: $0.30-5/M, 2.0 Flash: $0.40/M
❌ 需代理(Google 云服务在国内不可用) 免费额度

免费版:2.5 Pro/Flash 有免费 Rate Limit(15-30 RPM),超出按量计费

通义千问(阿里)

#5

Qwen (Alibaba)

输入 Qwen3.8-Max: ¥8/M, Qwen3.5-Max: ¥4/M, Qwen3.5-Plus: ¥2/M
输出 Qwen3.8-Max: ¥24/M, Qwen3.5-Max: ¥12/M, Qwen3.5-Plus: ¥6/M
✅ 国内直连 免费额度

新用户 100万 tokens 免费额度(90天有效);Qwen3.5 系列开源可自部署

Perplexity AI

#6

Perplexity AI

输入 Sonar Pro: $3/M, Sonar: $1/M
输出 Sonar Pro: $15/M, Sonar: $1/M
❌ 需代理 免费额度

API 无免费额度,Pro 订阅 $20/月(5000次搜索查询)

Meta Model API

#6

Meta Model API

输入 $1.25/1M tokens
输出 $4.25/1M tokens
❌ 需代理(Meta Model API 目前仅面向美国开发者公开预览,中国大陆需代理访问) 免费额度

Free tier: 60 RPM, 2M TPM limit; free credits on signup for US developers (public preview)

Exa

#7

Exa

输入 Pure pay-as-you-go by endpoint and search-type. Search $7/1k requests (base 10 results) + $1/1k per extra result + $1/1k AI summaries; Contents $1/1k pages; Answer $5/1k; Monitors $15/1k; Deep Search $12-15/1k; Agent fixed effort $0.012-$1.00/request or usage-based $0.10/ACU + tool calls (default $5 auto cap, $20 max cap). New accounts get $20 in free credits (~2,800 searches); Free Tier adds $10/month. No subscription, no minimum spend.
输出 Same as input. Enterprise custom volume + Zero Data Retention + SLA + postpaid invoice.
❌ Proxy required. Exa primary API runs at https://api.exa.ai from a single-region US deployment (San Francisco). No documented mainland China endpoint as of 2026-08-30. Mainland China production traffic typically needs a proxy or relay; transpacific first-byte latency is usually 100-300 ms. Enterprise plans can negotiate custom routing and dedicated infrastructure for regulated workloads (HIPAA / GDPR data residency). dashboard.exa.ai/onboarding provides agent-friendly onboarding that generates integration snippets tailored to the developer's stack, but API key issuance still requires stable international network access from China. 免费额度

Permanent $20 signup credit (~2,800 Search calls) + automatic $10 Free Tier monthly credit refill. No credit card, no minimum spend. Free credits apply to every endpoint.

Mistral AI

#9

Mistral AI

输入 Large 2: $2/M, Codestral: $1/M, Ministral 8B: $0.10/M
输出 Large 2: $6/M, Codestral: $3/M, Ministral 8B: $0.10/M
❌ 需代理 免费额度

注册送免费 API 额度(La Plateforme 每日速率限制内免费)

Cohere

#10

Cohere

输入 Command R7: $2.50/M, Command R: $0.50/M, Embed: $0.10/M
输出 Command R7: $10/M, Command R: $1.50/M
❌ 需代理 免费额度

免费 Trial API Key(2500次调用/月,限 Command R)

Fireworks AI

#11

Fireworks AI

输入 Firefunction-v2: $0.90/M, Llama 3.3 70B: $0.90/M, DeepSeek-R1: $2.00/M
输出 Firefunction-v2: $0.90/M, Llama 3.3 70B: $0.90/M, DeepSeek-R1: $2.00/M
❌ 需代理 免费额度

新用户 $0.50 免费额度

Black Forest Labs

#11

Black Forest Labs

输入 FLUX.2 [pro]: $0.05/MP, FLUX.1.1 [pro]: $0.04/MP, FLUX.1 [schnell]: $0.003/MP, FLUX.1 Fill: $0.05/MP
输出 Flat per-megapixel billing, no tier-based markup
❌ 需代理(BFL API 部署在 AWS US-East + EU-Frankfurt,国内访问需稳定代理) 免费额度

Free tier via API: limited generations per minute for [schnell] and [dev] endpoints; [pro] requires paid credits

Jina AI

#12

Jina AI

输入 Embeddings v3: $0.02/M tokens (10K free); Reranker v3: $0.018/M tokens (10K free); Reader: $0.02/M tokens (free up to 1M tokens); CLIP-v1: $0.02/M tokens
输出 Flat per-million-token billing, no tier-based markup; batch discount of 50% available on all endpoints
❌ 需代理(Jina API 部署在 AWS US-East + EU-Frankfurt,国内访问需稳定代理) 免费额度

10,000 free tokens/month on Embeddings + Reranker; Reader API free up to 1M tokens/month; no credit card required to start

Groq

#13

Groq

输入 Llama 3.3 70B: $0.59/M, DeepSeek-R1: $0.89/M, Mixtral 8x7B: $0.15/M
输出 Llama 3.3 70B: $0.79/M, DeepSeek-R1: $0.89/M, Mixtral 8x7B: $0.15/M
❌ 需代理 免费额度

免费速率限制(足够个人开发和小型测试使用)

Cerebras

#14

Cerebras

输入 $0.10/M to $0.60/M tokens(所有模型统一 $0.60/M input)
输出 $0.10/M to $0.60/M tokens(统一 $0.60/M output)
❌ 需代理 免费额度

免费速率限制(限制 RPM,适合测试)

AI21 Labs (Jamba)

#14

AI21 Labs (Jamba)

输入 Jamba 1.5 Mini: $0.20/M tokens, Jamba 1.5 Large: $2/M tokens
输出 Jamba 1.5 Mini: $0.20/M tokens, Jamba 1.5 Large: $2/M tokens
❌ 需代理 免费额度

无免费额度

CoreWeave

#14

CoreWeave

输入 GPU On-Demand hourly: NVIDIA HGX H100 $49.24, HGX H200 $50.44, HGX B200 $68.80, A100 80GB $21.60, L40S $18.00, L40 $10.00, GB200 NVL72 $42.00. Spot discounts: H100 $19.71, H200 $20.93, B200 $34.11, A100 $9.65. Serverless Inference is pay-per-token via W&B Inference; Dedicated Inference is GPU-hour / node-based.
输出
❌ 无中国大陆直连节点。数据中心集中在美国;2026 年已扩展至印尼(首个 APAC 节点)、英国(两个数据中心运营中)、瑞典。国内访问需代理,跨太平洋延迟较高。 免费额度

无永久免费层。Serverless Inference 按 token 计费;GPU 实例按小时计费(On-Demand)或有 spot 折扣。0EM(零出口费迁移)计划在迁移阶段免除出口费。

DeepInfra

#15

DeepInfra

输入 Llama 3.3 70B: $0.49/M, DeepSeek-V3: $0.99/M, DeepSeek-R1: $1.09/M
输出 Llama 3.3 70B: $0.73/M, DeepSeek-V3: $0.99/M, DeepSeek-R1: $1.09/M
❌ 需代理 免费额度

注册即送免费额度(每日有限免费调用)

NVIDIA NIM

#15

NVIDIA NIM

输入 Partner-dependent (Nemotron-3 Ultra 550B: $0.50-0.90/1M tokens)
输出 Partner-dependent (Nemotron-3 Ultra 550B: $1.70-3.60/1M tokens)
❌ 需代理(build.nvidia.com 平台为美国服务,受出口管制限制) 免费额度

免费开发端点(无需信用卡,生成 API Key 即可开始)

Amazon Bedrock

#15

Amazon Bedrock

输入 Claude Sonnet 4.5: $6/M, DeepSeek V3.2: $0.62/M, Mistral Large 3: $0.50/M, GPT-5.5: $5.50/M, GPT-5.4: $2.75/M, Grok 4.3: $1.25/M, Qwen3 235B: $0.23/M, Kimi K2.5: $0.60/M
输出 Claude Sonnet 4.5: $30/M, DeepSeek V3.2: $1.85/M, Mistral Large 3: $1.50/M, GPT-5.5: $33/M, GPT-5.4: $16.50/M, Grok 4.3: $2.50/M, Qwen3 235B: $0.91/M, Kimi K2.5: $3/M
❌ 大陆用户需代理(AWS 中国区(由光环新网/西云数据运营)尚未提供 Bedrock 服务,需使用全球区域 + VPN/代理) 免费额度

AWS Free Tier 赠送新用户最高 $200 抵扣金(6个月有效)。Bedrock 按用量抵扣,无永久免费额度

Hyperbolic

#15

Hyperbolic

输入 Per-token inference scaled from market GPU rates; H100 from $2.89/GPU-hour, H200 $3.49/GPU-hour
输出 Pay-as-you-go GPU compute billed per hour; reserved discounts for committed capacity
❌ 需代理 免费额度

无长期合同;按量计费按小时结算;无永久免费档,但无配额上限、即开即用

Fireworks AI

#15

Fireworks AI

输入 Serverless (Standard, $/1M tokens): DeepSeek V4 Pro $1.74、V4 Flash $0.22、Kimi K3 $3.00、K2.7 Code $0.95、MiniMax M3 $0.30、Qwen 3.7 Plus $0.40、GPT OSS 120B $0.15、GLM 5.1 $1.40。嵌入按模型档 $0.008-$0.016/1M。按需 GPU 按小时:H100 $7.00、H200 $7.00、B200 $10.00、B300 $12.00、GB300 $18.00(9 月 1 日上调)。
输出 Serverless 输出($/1M):DeepSeek V4 Pro $3.48、V4 Flash $0.66、Kimi K3 $15.00、MiniMax M3 $1.20、Qwen 3.7 Plus $1.60、GPT OSS 120B $0.60、GLM 5.1 $4.40。缓存输入显著更低(如 DeepSeek V4 Pro $0.145)。训练:SFT 每 1M 训练 token $0.50(≤16B)到 $10.00(>300B);Serverless Training API 仅按 token 计费、无闲置 GPU 成本。
❌ 无中国大陆直连端点。作为美国平台,大陆直连需代理或海外中转,跨太平洋首字节延迟通常 150-250ms。最近 APAC 端点为东京多区域。 免费额度

注册即享 $1 免费 serverless 额度(pay-per-token,postpaid billing);无永久免费档。按需 GPU 与训练按用量计费。

SambaNova

#16

SambaNova

输入 Llama 3.3 70B: $0.50/M, DeepSeek-R1: $0.80/M, Qwen 72B: $0.40/M
输出 Llama 3.3 70B: $0.50/M, DeepSeek-R1: $0.80/M, Qwen 72B: $0.40/M
❌ 需代理 免费额度

免费速率限制(测试用途)

Nebius

#16

Nebius

输入 Per-token OpenAI-compatible inference: DeepSeek-V4-Flash $0.14/M input, $0.28/M output; Kimi K3 $3/$15; MiniMax M3 $0.30/$1.20; Llama 3.3 70B $0.13/$0.40
输出 Two flavors per model - Base and Fast (-fast suffix) - same outputs, Fast trades higher token price for lower latency via speculative decoding
❌ 需代理 免费额度

注册即可起步,无月费;按量付费(pay-as-you-go)按 token 计费;弹性动态限流,随用量自动扩容至基础额度 20 倍,无需预购容量

Anyscale

#17

Anyscale

输入 Llama 3.3 70B: $0.39/M, DeepSeek-R1: $1/M, Mixtral 8x22B: $0.90/M
输出 Llama 3.3 70B: $0.39/M, DeepSeek-R1: $1/M, Mixtral 8x22B: $0.90/M
❌ 需代理 免费额度

新用户 $100 免费额度(一次性),另提供少量项目启动额度($3-$5)

DigitalOcean Gradient

#17

DigitalOcean Gradient

输入 Llama 3.3 70B: $0.50/M, DeepSeek-R1: $0.80/M
输出 Llama 3.3 70B: $0.50/M, DeepSeek-R1: $0.80/M
❌ 需代理 免费额度

新用户 $100 免费额度(60 天有效)

Stability AI

#18

Stability AI

输入 SD3.5: $0.0065/image, SDXL: $0.01/image, Stable Code: $0.50/M tokens
输出 按图像/视频/音频输出计费,非 token 计费
❌ 需代理 免费额度

新用户 25 次免费生成

Lepton AI

#19

Lepton AI

输入 Llama 3.3 70B: $0.80/M; Llama 3.1 405B: $3.50/M; Qwen 2.5 72B: $0.80/M; DeepSeek-R1: $2.00/M
输出 Same rate as input for most models (symmetric pricing)
⚠️ 部分可用(ap-northeast-1 东京 region 距离最近,国内访问需代理) 免费额度

$5 试用额度(30 天有效,足够完整评估任意一款主机模型)

可灵 AI(快手)

#19

Kling AI (Kuaishou)

输入 Per-second video billing. Kling 3.0 Turbo: 720p $0.112/s, 1080p $0.14/s. Kling 3.0: 720p $0.084-0.126/s, 1080p $0.112-0.168/s (Native Audio), 4K $0.42/s. Kling 3.0 Omni: 720p $0.084-0.126/s, 1080p $0.112-0.168/s, 4K $0.42/s, varies by video/audio input.
输出 Image API: Kling Image 3.0 $0.028/image (1K/2K), 3.0-omni $0.028-0.056/image (up to 4K), Image 2.1 $0.014-0.028/image, Image O1 $0.028/image, Multi-Shot $0.07/call. Video list price 1 Unit = $0.14; image list price 1 Unit = $0.0035.
✅ 官方支持部分中国区域调用(快手自研,国内云节点可用;海外 klingai.com 需网络环境) 免费额度

新用户注册后可获得少量免费体验积分(平台内赠送 credits),可体验部分视频/图片生成;正式 API 使用无永久免费套餐,按生成的秒数/张数计费。

ElevenLabs

#20

ElevenLabs

输入 TTS: $0.03-0.30/1K chars depending on model/quality; STT: $0.47/hour; Sound Effects: $0.04/request
输出 Same as input rate per model
⚠️ 部分可用(需稳定代理) 免费额度

10,000 credits/month (≈10 minutes audio, free plan with limited features)

Ideogram

#20

Ideogram

输入 Ideogram 3.0: $0.04/image (standard), $0.08/image (Turbo)
输出 4 images per generation (default), additional images via paid credits
❌ 需代理 免费额度

新用户赠送 $1 体验金,约 25 次标准生成

Lambda GPU Cloud

#20

Lambda

输入 GPU instance on-demand per GPU/hr: NVIDIA B200 SXM6 $6.69, H100 SXM $3.99, A100 SXM 80GB $2.79, A100 SXM 40GB $1.99, GH200 $2.29, A100 PCIe $1.99, A6000 $1.09, A10 $1.29, V100 $0.79. 1-Click Cluster reserved per GPU/hr: B200 $9.86 (16 GPU) to $8.87 (256+), H100 $6.16 to $5.54. Superclusters / Private Cloud (4000+ GPUs) via sales. No egress fees.
输出 Billed by the second per GPU instance (no idle GPU cost); clusters and reserved capacity carry a minimum commitment (2 weeks to 1 year). Managed orchestration: Managed Kubernetes and Slurm included at no extra per-node fee; Lambda Stack one-line install. No per-token inference pricing — Lambda's Inference API is winding down; self-hosted inference billed as GPU runtime.
❌ 无中国大陆直连节点;数据中心在美国(加州等区域),国内访问需代理或海外中转。 免费额度

无永久免费 GPU 层;注册后按用量计费(GPU 实例按秒计费)。1-Click Clusters 与 Superclusters 需预付 / 承诺周期(2 周至 1 年)。

Deepgram

#21

Deepgram

输入 STT: $0.0043-0.012/min; TTS: $0.015-0.030/1K chars
输出 N/A
⚠️ 部分可用(需代理,无中国大陆数据中心) 免费额度

New users get $200 in free credits; no credit card required

RunPod

#21

RunPod

输入 Per-second GPU: H100 SXM $2.99/hr, A100 SXM $1.49/hr, B200 $5.89/hr, L40S $0.99/hr, RTX 4090 $0.69/hr
输出 Same as input (per-GPU-second — no idle fees, no egress fees)
⚠️ 部分可用(基础设施在全球 31 区域,国内直连不稳定,需稳定代理) 免费额度

$5 starter credit on signup, no credit card required; per-second billing means spiky workloads never pay for idle

Luma AI

#21

Luma AI

输入 Credit-based subscription. Plans: Plus $30/mo (10,000 credits), Pro $90/mo (40,000 credits), Ultra $300/mo (150,000 credits). Ray 3.2 video: 1080p text-to-video 400 credits/5s, 720p 100 credits/5s, Draft 20 credits/5s; Seedance 2.0 1080p 240 credits/sec, 4K 959 credits/sec.
输出 Image credits: Uni-1 30 credits/image, Seedream 1-3 credits/image, GPT Image 2 from 3 (Low-1K) to 255 (High-4K) credits. Audio: ElevenLabs v3 TTS 21 credits/1,000 chars. Utilities: background removal 1 credit/image, reframe video 32 credits/sec, upscale to 4K 17 credits/sec.
❌ 需代理(无中国大陆直连节点,API 托管于全球 AWS/GCP 边缘) 免费额度

无免费 API 套餐。Luma 提供有限免费信用额度供新用户体验(平台注册后少量免费 credits),但正式 API 使用需至少订阅 Plus $30/月(10,000 credits)。

Voyage AI

#22

Voyage AI

输入 voyage-4-large: $0.12/M, voyage-4: $0.06/M, voyage-4-lite: $0.02/M, voyage-context-3: $0.18/M, voyage-code-3: $0.18/M, voyage-multimodal-3.5: $0.12/M
输出 (embedding API — no output tokens; rerank-2.5: $0.05/M tokens, rerank-2.5-lite: $0.02/M tokens)
⚠️ 部分可用(官网在中国可访问,但 API 在中国大陆的稳定性未公开声明) 免费额度

200M free tokens per account (text embedding + reranker), 50M free tokens for legacy domain models, 200M free text tokens + 150B free pixels for multimodal

Runway

#22

Runway

输入 Gen-4 Turbo: 50 credits/5s clip (~$0.50), Gen-4 Alpha: 100 credits/5s (~$1.00), Act-Two: 150 credits/5s (~$1.50), Frames: 200 credits/sequence (~$2.00)
输出 Credit-based system (~$0.01/credit). Credit packs: $12 (1,000 credits) to $200 (25,000 credits). Enterprise: custom pricing.
❌ 需代理(AWS US-East + GCP US-Central 托管) 免费额度

无免费 API 套餐。最低 $12 购买 1,000 credits,约 20 次 Gen-4 Turbo 生成。Web 应用免费但无 API 访问。

TwelveLabs

#22

TwelveLabs

输入 Free plan: 600 minutes video indexing + API usage (Pegasus 1.2 $0.042/min indexing one-time, $0.021/min input video, $0.0075/1k tokens output text, $0.0015/min monthly embedding infrastructure). Developer plan: paid overage, 90-day index retention lifted to unlimited.
输出 Same per-minute and per-1k-tokens billing. Marengo embeddings billed separately by indexed minute.
❌ 需代理(无中国大陆直连节点,API 托管于 AWS us-west-2 / us-east-1 区域) 免费额度

免费计划:注册即获 600 分钟视频索引额度(一次性累积额度),可调用全部核心 API(Pegasus 1.2/1.5、Marengo 3.0、Search、Embed、Analyze & Segment)。索引数据保留 90 天;升级 Developer 计划后索引保留变为无限。

Mixedbread AI

#22

Mixedbread AI

输入 按量计费,三块:索引(Fast $1.50/百万内容 token;High Quality 含 OCR/转写/摘要/多模态增强 $3/百万)、搜索(Grep $0.10/千次、Semantic $4/千次、Toast 1 $1/千次,重排 +$3.50 或 +$1.50/千次)、存储 $0.50/百万内容 token/月。Toast 1 按 LLM token:输入 $0.50/百万(发布价 $0.30)。
输出 Toast 1 输出 $1.20/百万 LLM token(发布价 $0.72,40% 折扣);缓存输入 $0.06/百万(发布价 $0.036),缓存写入免费。Semantic 检索重排加价 $3.50/千次。
❌ 无中国大陆直连端点。Mixedbread 是美国/欧洲平台(法人实体 mixedbread ai inc.,柏林/旧金山),大陆直连需代理或海外中转,跨太平洋首字节延迟通常 150-250ms。 免费额度

Starter 免费送 $5 一次性额度,无需信用卡;3 个工作区用户、10 个 Store、每分钟 100 次请求,适合评估真实语料。无永久免费档。

Crusoe Cloud

#23

Crusoe

输入 Pay-as-you-go per 1M tokens, four tiers by parameter count: <16B ($0.40) / 16B-70B ($2.50) / 70B-300B ($6.00) / >300B ($10.00). Serverless Fine-Tuning follows identical 4-tier structure. Managed Inference spans DeepSeek V3/V4 Pro/V4 Flash, GLM 5.1/5.2, GPT-OSS 20B/120B, Gemma 4 31B-it, Kimi K2.6, Llama 3.1/3.3, Nemotron 3 family (VoiceChat/Ultra/Lightning/Nano/Super/Nano Omni), Qwen3 family (8B/235B/Qwen3.5/Qwen3.6), Yutori n1.5. Cached tokens at $0.03-$1.50 per 1M.
输出 Self-Serve Deployments: NVIDIA H100 80GB HGX $5.50/hr, NVIDIA H200 141GB HGX $6.00/hr (dedicated endpoints for open + fine-tuned models). Tailored Deployments and Provisioned Throughput via sales engagement. Managed Kubernetes $0.10/cluster hour; Container Registry $0.10/GiB-month; Object Storage $0.06/GiB-month.
❌ 无中国大陆直连节点。托管推理 API 托管于 Crusoe 美国数据中心(Colorado / Texas 等);国内访问需代理。 免费额度

无永久免费层;按使用量计费;可联系销售获取批量折扣。Serverless Fine-Tuning 四档起价 $0.40/M tokens。

Letta

#23

Letta

输入 按用量:Free $0(最多 3 个有状态 agent、BYOK)、Pro $20/月(最多 20 个 agent、Letta Auto 周/月额度 + 按量超额)、Teams Pro 按席位(共享 agent、权限)、Developer Plan 纯 API 调用按信用(无 agent 上限)、Enterprise 商务报价(自托管/定制模型)。服务端工具按 CPU 时间 $0.00015/秒计费。
输出 Letta Auto 用量超出 Pro 套餐限额后按 pay-as-you-go 模型价(参考 platform.letta.com/models)。Mods、Skills、Conversations 等功能均包含在套餐内,不单独计费。远程 MCP 工具由 MCP 供应商执行,不消耗 Letta 信用。
❌ 无中国大陆直连端点。Letta, Inc. 注册于美国旧金山,platform.letta.com 是单一全球端点;大陆生产流量需要代理或中转,跨太平洋首字节延迟通常 150-250ms。Agent harness 完全开源(Apache-2.0),可自托管在自有基础设施上规避直连限制。 免费额度

Free $0/月(最多 3 个有状态 agent),自带 API key;Letta Auto 用量与 Skills 受限,可用于评估和小型个人助手。

Writer

#25

Writer

输入 Palmyra X5: $0.60/M, Palmyra X4: $2.50/M
输出 Palmyra X5: $6.00/M, Palmyra X4: $10.00/M
⚠️ 部分可用(官网在中国可直接访问,但 API 稳定性未公开声明) 免费额度

14-day free trial, no credit card required (Starter plan); no standalone free API credits

Reka

#25

Reka

输入 Reka Edge $0.10/M 输入;Reka Flash $0.80/M;Reka Core $2.00/M;Reka Flash Research 按请求计费($25-$60/千次)。按量付费,无最低承诺。
输出 Reka Edge $0.10/M;Reka Flash $2.00/M;Reka Core $6.00/M。多模态(图像、视频、音频)单独按"每分钟"计费,Flash 模型视频 $0.06/分钟、Core $0.08/分钟;图像按调用计费(Flash $0.01、Core $0.02)。
❌ 无中国大陆直连端点。Reka 注册于美国加州 Sunnyvale(530 Lawrence Expressway),api.reka.ai 是单一全球端点;大陆生产流量需代理或中转,跨太平洋首字节延迟通常 150-250ms。开源权重可自托管(NVIDIA GPU ≥24GB VRAM),规避直连限制。 免费额度

通过 platform.reka.ai 注册即可获取 API key;Pay-as-you-go 用量计费,不强制最低承诺,可在充值额度内自由试用。无永久免费档、无月赠送额度。

Aleph Alpha

#28

Aleph Alpha

输入 Enterprise contract — no public per-token price (quote-based; PhariaAI on-prem deployment custom-priced)
输出 Same — enterprise quote, no public per-token list
❌ 不适用(德国主权 AI 厂商,受欧盟出口管制约束) 免费额度

No public free tier — self-serve signup creates an account, but pricing is enterprise contract only

Cartesia

#28

Cartesia

输入 Sonic: $0.03/1K chars (Standard), $0.06/1K chars (Pro), $0.12/1K chars (Turbo); Sonic 2: $0.04-0.15/1K chars
输出 TTS character-based pricing (output = audio), per-second for streaming
⚠️ 部分可用(需稳定代理) 免费额度

10,000 characters/month free tier (≈10 minutes), no credit card required

Hume AI

#28

Hume AI

输入 EVI 3: ~$0.096/min (pay-per-second); OCTAVE TTS: $0.048/1K chars; Expression Measurement: $0.0008/sec
输出 Same as input (most endpoints symmetric)
Proxy required (US export controls) 免费额度

Free credits on signup (amount per account), enough to trial EVI 3

Modal

#28

Modal

输入 按 GPU 秒计费:H100 $0.001097/s、A100 80GB $0.000694/s、T4 $0.000164/s
输出 同输入(按 GPU 时间,无 idle 费用)
❌ 需代理(基础设施在 AWS/GCP 海外区域,国内直连不稳定) 免费额度

Starter 计划免费 + $30/月 计算额度(每月刷新)

AssemblyAI

#30

AssemblyAI

输入 按音频小时计费:Universal-3.5 Pro $0.21/hr、Universal-2 $0.15/hr(pre-recorded)
输出 N/A(speech-to-text 计费基于音频时长,无 token 概念)
❌ 需代理(基础设施在 US,国内直连不稳定) 免费额度

免费账户每月 185 小时 pre-recorded + 333 小时 streaming,无需信用卡

Together AI

#39

Together AI

输入 $0.03-$1.74/MTok(模型相关)
输出 $0.12-$4.50/MTok(模型相关)
⚠️ 需代理(美国服务商) 免费额度

注册赠送 $1 初始额度(需绑定支付方式)

FriendliAI

#

FriendliAI

输入 GLM-5.2/5.1: $1.4/M, DeepSeek-V3.2: $0.5/M, Qwen3-235B: $0.2/M, MiniMax-M2.5: $0.3/M, Gemma-4-31B: $0.14/M
输出 GLM-5.2/5.1: $4.4/M, DeepSeek-V3.2: $1.5/M, Qwen3-235B: $0.8/M, MiniMax-M2.5: $1.2/M, Gemma-4-31B: $0.4/M
⚠️ Partially available — servers in US & Korea, ~200-400ms latency from CN 免费额度

No free tier — per-token billing with no minimum

Suno

#40

Suno

输入 Free: $0 (10 songs/day, non-commercial); Pro: $10/mo (500 songs, commercial rights); Premier: $30/mo (2,000 songs, Studio)
输出 Annual: Pro $96/yr (20% off), Premier $288/yr (20% off); credit top-ups available as add-ons
⚠️ 国内访问受限(需稳定海外代理),无官方中国大陆节点 免费额度

免费版每天 10 首歌曲(不可商用),新账户直接开通,无需信用卡

Liquid AI

#40

Liquid AI

输入 无按 token 云 API 定价——LFM 权重免费下载/运行/微调(LFM Open License,公司年收入超过 1000 万美元前免版税)。推理要么在自有硬件上自托管,要么通过 LEAP SDK 本地部署;无 per-token 请求费用。
输出 无 per-token 输出计费。模型在 CPU/GPU/NPU 上本地或自托管运行,适合对延迟与隐私敏感的高频或离线工作负载(Liquid 的定位:「为何为云 API 按 token 付费,而非本地推理」)。
⚠️ 无托管云 API,因此无中国直连端点问题。模型权重自由下载并可在本地/自托管运行,包括在中国境内(对数据主权友好)。不需要云 API 网关或代理;token 处理完全在本地硬件上进行。 免费额度

全部 14+ 个 LFM 模型在 LFM Open License 下免费下载/运行/微调——包括商用(公司年收入 ≤ 1000 万美元)。研究、教育与非营利使用永久免费、无收入上限。

Nscale

#

Nscale

输入 Nscale 采用充值余额(credit)按量计费,价格按模型分列,在 Nscale 控制台 AI Services → Models 与 /v1/models API 中按每 1M 输入 token 报价(文本模型);图像模型输入按 0 计费。定价结构经 serverless-openapi.yaml ModelPricing schema 验证(input = 每 1M 输入 token 美元价)。
输出 输出按每 1M 输出 token 计费(文本模型);图像模型按每百万像素计费(ModelPricing.output 定义为 per-million output pixels)。具体美元单价在控制台与 /v1/models 端点内按模型展示,注册充值后可查。
❌ 无中国大陆直连端点。inference.api.nscale.com 是单一全球端点,数据中心集中于欧洲(英国 Loughton、挪威 Glomfjord/Narvik)与美国(德州、西弗吉尼亚);大陆生产流量需代理或中转,跨大西洋/跨太平洋首字节延迟通常 200-350ms。主打 EU 数据主权与在国境内处理,适合欧洲合规场景而非中国区低延迟。 免费额度

无永久免费档。新用户注册后通过 Google SSO 登录,官方提示早期用户可能获得免费信用额度(promotional offer,非保证)。默认需先充值(credit)才能调用 serverless 推理 API。

Morph

#41

Morph

输入 按模型分别计费,按每 1M 输入 token 计(文本模型)。从 llms.txt 验证(2026-08-28):Kimi K3 $2.90、Qwen 3.5 397B $0.50、GLM-5.2 744B $1.10、GLM-5.3-Flash $0.15、MiniMax M3 $0.30、MiniMax M2.7 $0.279、DeepSeek V4 Flash (beta) $0.12、Qwen 3.8/3.6 27B $0.289、Gemma 4 31B $0.14。Qwen 3.5 397B 缓存输入 $0.30 per 1M(其他模型无独立缓存价)。Fast Apply 用 `auto` 模型按 prompt+completion 两档计(OpenRouter 验证 morph-v3-fast $0.0008/$0.0012 per 1k、morph-v3-large $0.0009/$0.0019 per 1k),Compact、Reflexes、WarpGrep 与 Router 按事件计费(详见 pricingEN)。
输出 按模型分别计费,按每 1M 输出 token 计(文本模型)。从 llms.txt 验证:Kimi K3 $14.00、Qwen 3.5 397B $3.50、GLM-5.2 744B $4.10、GLM-5.3-Flash $0.50、MiniMax M3 $1.20、MiniMax M2.7 $1.20、DeepSeek V4 Flash (beta) $0.278、Qwen 3.8/3.6 27B $2.40、Gemma 4 31B $0.40。Fast Apply 输出按 morph-v3-fast $1.20 / morph-v3-large $1.90 per 1M token;Fast Apply 用 `auto` 推荐档(5,000-10,500 tok/s, ~98% 精度)按实际路由模型分档计费。Kimi K3 `service_tier: standby` 走 GLM-5.2 同等待遇,标准按 token 计。
❌ 无中国大陆直连端点。base URL https://api.morphllm.com 为单一全球端点,Y Combinator 背书的美国总部(AutoInfra, Inc.);OpenRouter 上 morph/morph-v3-fast 和 morph/morph-v3-large 走多区域路由,但官方端点不在大陆。中国大陆生产流量需代理或中转,跨太平洋首字节延迟通常 200-400 ms。 免费额度

无永久免费档(无零元 starter plan)。新用户注册即可获得 API key 并按 pay-as-you-go 充值;官方为初创公司提供 Startup Credits(最高 $5,000),但需单独联系销售申请。Morph 主页明确「Start free. Get API Key」,但实际计费从首次 token 调用开始。

Firecrawl

#42

Firecrawl

输入 纯 credit-based 按 endpoint 与特性消耗,每月按 plan 重置额度;额外 credit 按 $5 一档购买(每档 1k–5k credits 不等)。从 firecrawl.dev/pricing 与 docs.firecrawl.dev/billing.md 验证(2026-08-29):Free 1k/月 $0;Hobby 5k/月 $19(年付 $16);Standard 100k/月 $83(年付 $99);Growth 500k/月 $333(年付 $399);Scale 1M/月 $599;Enterprise Custom。Credit 单价:Scrape / Crawl / Map 1 cr/page、Search 2 cr/10 results、Monitor 1 cr/page/check、Interact 2–7 cr/browser-minute。Batch Scrape / Extract 按子操作同价。Agent 5 daily runs free,超量动态计费。
输出 无独立 output 维度——Firecrawl 输出的是 scraped markdown / HTML / JSON 内容,计费仅按输入端的 credit 消耗与并发浏览器占用。出站数据本身不另收费;唯一例外是 Search 在 ZDR(zero-data-retention)企业档下按企业价另议。
❌ 无中国大陆直连端点。Firecrawl 总部位于美国旧金山(YC W23),API base URL https://api.firecrawl.dev 单区域部署;docs.firecrawl.dev/billing.md 未列出任何亚太 / 中国路由层。中国大陆生产流量需代理或中转,跨太平洋首字节延迟通常 200-400 ms;Free 档 1000 credits 评估期可代理试用,但稳定流量建议自建代理或自部署 Firecrawl 自托管版(self-hosted Docker / Kubernetes)。 免费额度

Free 档永久免费:1,000 credits / 月,无需信用卡,支持 Scrape / Crawl / Map / Search / Batch Scrape / JSON mode 全部 endpoints;2 concurrent browsers;50,000 max queued jobs。适合评估与轻量使用,但 credit 重置周期每月初清零、当月用完不再补。Agent 5 daily runs free 适用于所有 plan(含 Free)。

🔗

AI API 网关与聚合器

一个 API Key → 400+ 模型。统一访问、自动路由、灵活切换模型。

OpenRouter

#3

300+ 模型

核心优势

300+ 模型,覆盖所有主流厂商

❌ 需代理(未封锁,但国内访问不稳定) 免费额度

无初始免费额度,免信用卡注册

Vercel AI Gateway

#8

200+ 模型

核心优势

$5/月免费 Credit,Token 零加价(含 BYOK)

⚠️ 需代理(境内需 VPN) 免费额度

每团队每月 $5 免费 AI Gateway Credits(仅限免费层可用模型)

Dify

#9

100+ 模型

核心优势

GitHub 150,010 stars 的 LLM 应用开发平台事实标准(langgenius/dify,Apache 风格许可)

✅ 国内母公司 LangGenius,原生中文支持、阿里云一键部署、SOC 2/GDPR/ISO 27001 认证,国内云完全可用 免费额度

Sandbox 免费层:200 GPT-4 调用 + 5MB vector 存储 + 10 文档 + 2 工作流 + 5,000 API 调用/月(30 天日志保留)

Cloudflare AI Gateway

#10

500+ 模型

核心优势

统一的 API 网关:为 20+ AI API Provider 提供单入口

✅ 全球 CDN(国内边缘节点通过合作伙伴运营) 免费额度

每月 100 万次请求免费(Gateway 服务费)

Portkey

#11

200+ 模型

核心优势

🎯 AI Gateway + 全栈可观测性一体化:Gateway、Logs、Feedback、Guardrails 同一控制台

❌ 需代理(portkey.ai 域名在国内访问不稳定,推荐自托管开源版本) 免费额度

每月 10,000 次请求;100K 日志保留;社区 Slack 支持;BYOK 免费;无信用卡要求

LiteLLM

#12

100+ 模型

核心优势

100+ Provider 统一接口的知名开源方案

✅ 开源自部署,无限制 免费额度

开源版完全免费(自托管)

硅基流动

#12

100+ 模型

核心优势

国内一站式模型托管,100+ 开源模型在线

✅ 国内直连 免费额度

注册送 ¥1-200 免费额度(按活动调整),足够跑 1-20M tokens

Arize Phoenix

#12

100+ 模型

核心优势

唯一的 Elastic-2.0 开源(10,600+ 颗星),OpenTelemetry 原生 LLM 可观测性平台

✅ 开源自部署 + Cloud 托管(境外服务,境内需代理访问 SaaS) 免费额度

Phoenix Cloud 免费套餐:每工作空间 10 GiB 存储(无时间限制)

Block AI(b.ai)

#13

150+ 模型

核心优势

一个 Key 访问 GPT-4o / Claude / Gemini 等海外模型

✅ 国内直连 免费额度

注册送体验额度(每天免费请求次数待确认)

fal.ai

#13

200+ 模型

核心优势

200+ 模型一站式:FLUX、Kling、Hunyuan、MiniMax、Wan 全部接入

❌ 需代理(fal.ai 部署在 AWS US-East,国内访问需稳定代理) 免费额度

$1 free credit on signup (no credit card required); 100+ models available free for first-time evaluation

Helicone

#13

100+ 模型

核心优势

YC W23 出品,Apache-2.0 开源 LLM Observability 平台

✅ 开源自部署 + SaaS 免费额度

每月 10,000 次免费请求 + 1GB 存储 (Hobby)

FreeModel

#14

50+ 模型

核心优势

DeepSeek 官方合作商

✅ 国内可用(直连) 免费额度

注册送额度(具体金额待确认)

APIKEY.FUN

#15

50+ 模型

核心优势

国内直连访问海外模型

✅ 国内直连 免费额度

注册送少量免费额度

Replicate

#16

50000+ 模型

核心优势

50,000+ 开源模型,覆盖 AI 全部领域

❌ 需代理 免费额度

新用户 $0.50 免费额度

Novita AI

#18

50+ 模型

核心优势

Llama 3.3 70B 定价有竞争力

⚠️ 部分可用(新加坡节点,国内延迟尚可) 免费额度

$0.50 免费试用额度

Baseten

#18

35+ 模型

核心优势

Per-second GPU 计量,与 Replicate 同类竞品中按量计费最透明

⚠️ 部分可用(需稳定代理连接) 免费额度

$30 in inference credits on signup (valid 30 days); custom model hosting free tier: 1 deployed model

Mem0

#18

50+ 模型

核心优势

AI Agent 记忆层的事实标准(61,493 GitHub stars, Apache-2.0,mem0ai/mem0)

✅ 开源(Apache-2.0,61k+ stars)+ Cloud SaaS;国内自部署无障碍 免费额度

Hobby 免费层:10,000 memories + 1,000 retrieval API calls/月(无限 end users),社区支持

Tavily

#18

7+ 模型

核心优势

🎯 AI agent 默认搜索 API:LangChain TavilySearchResults、LlamaIndex TavilyToolSpec、AutoGen/CrewAI/smolagents 全部默认集成

⚠️ Tavily API 部署在 AWS US-East / EU-West;国内直连延迟 200-400ms;推荐用 Cloudflare Worker 或腾讯云 Edge Function 中转(降至 50-100ms);无 CN 节点计划;无国内信用卡支付渠道 免费额度

Free 永久:1,000 credits/month(滚动 30 天窗口),无信用卡、无功能限制;Search/Extract/Crawl/Map/AI Extraction 各占独立配额(1,000 Search + 500 Crawl);超额返 HTTP 429,无超额计费

Hugging Face

#20

200000+ 模型

核心优势

全球最大 AI 模型社区,200,000+ 模型

❌ 需代理(huggingface.co 在国内被墙) 免费额度

免费推理 API(有限速率限制)

fal.ai

#27

1396+ 模型

核心优势

🧠 1,396+ 开源/商业模型:FLUX.2 Pro、GPT Image 2、Kling Video、Nano Banana 2、ElevenLabs TTS、MiniMax 等

✅ 大陆直连可用(fal.ai 使用全球 CDN,无需代理);延迟因模型推理节点而异,典型 200-500ms 免费额度

无免费 API 层(预付 $10 起买信用余额);无月费、无订阅、无隐藏费用;Serverless 模式下零运行时自动缩零

💰 Token 价格对比表

厂商 热门模型 输入 输出 国内可用 免费额度
OpenAI GPT-5, GPT-5-mini GPT-5: $1.25/M, GPT-5-mini: $0.25/M, GPT-4o: $2.50/M, GPT-4o-mini: $0.15/M GPT-5: $10/M, GPT-5-mini: $2/M, GPT-4o: $10/M, GPT-4o-mini: $0.60/M ❌ 需代理(受出口管制限制,国内用户需使用代理访问) 新账户赠 $5 额度,有效期 3 个月 查看详情 →
Anthropic Claude Claude Sonnet 5, Claude Opus 5 Sonnet 5 (intro): $2/M, Opus 5: $5/M, Sonnet 5/4.5: $3/M, Opus 4.8: $5/M, Haiku 4.5: $0.80/M Sonnet 5 (intro): $10/M, Opus 5: $25/M, Sonnet 5/4.5: $15/M, Opus 4.8: $25/M, Haiku 4.5: $4/M ❌ 需代理(受美国出口管制,国内用户需代理) Claude 免费版有限额度(网页端);API 端无赠送额度;Sonnet 5 限时价 $2/M(至 2026-08-31) 查看详情 →
Azure OpenAI(微软 Azure) GPT-4o, GPT-4o-mini GPT-4o: $2.50/M, GPT-4o-mini: $0.15/M, o1: $15/M, o3: $10/M GPT-4o: $10/M, GPT-4o-mini: $0.60/M, o1: $60/M, o3: $40/M ⚠️ 世纪互联运营版(国内可用,模型比国际版少) 新用户 $200 免费额度(30天有限) 查看详情 →
OpenRouter GPT-4o, Claude Opus/Sonnet/Haiku $0.065/M to $15/M 取决于模型,平均 $0.50/M $0.26/M to $75/M 取决于模型,平均 $2/M ❌ 需代理(未封锁,但国内访问不稳定) 无初始免费额度,免信用卡注册 查看详情 →
DeepSeek(深度求索) DeepSeek-V3, DeepSeek-R1 DeepSeek-V3: ¥0.14/M token (≈$0.02), R1: ¥0.14/M DeepSeek-V3: ¥0.28/M token (≈$0.04), R1: ¥0.28/M ✅ 国内直连 注册送 ¥500 免费额度(约500万tokens) 查看详情 →
xAI Grok Grok 4.6, Grok 4.6 Fast Grok 4.6: $2/M tokens (fast variant $4/M) Grok 4.6: $6/M tokens (fast variant $12/M) ❌ 需代理 无免费 API 额度(Grok Build 试用;初期免信用卡) 查看详情 →
Gemini 多模态 Gemini Omni Flash (multimodal, 1M context, audio+video+text), Nano Banana 2 Lite (image generation, $0.034/1K tokens) Omni Flash: $0.15/M input tokens (multimodal); Nano Banana 2 Lite: $0.034/1K image tokens; Veo 3.1: $0.10/sec video Omni Flash: $0.60/M output tokens (multimodal); Nano Banana Pro: $0.12/1K image tokens ❌ 需代理(Google AI Studio 在中国大陆不可直接访问,需稳定代理) Free tier: Gemini Omni Flash 15 RPM, Veo 3.1 2 generations/day, Imagen 4 10 images/day — no credit card required 查看详情 →
Google Gemini Gemini 2.5 Pro, Gemini 2.5 Flash 2.5 Pro: $1.25-10/M, 2.5 Flash: $0.15-1.25/M, 2.0 Flash: $0.10/M 2.5 Pro: $5-40/M, 2.5 Flash: $0.30-5/M, 2.0 Flash: $0.40/M ❌ 需代理(Google 云服务在国内不可用) 免费版:2.5 Pro/Flash 有免费 Rate Limit(15-30 RPM),超出按量计费 查看详情 →
通义千问(阿里) Qwen3.8-Flash-Next, Qwen3.8-Flash Qwen3.8-Max: ¥8/M, Qwen3.5-Max: ¥4/M, Qwen3.5-Plus: ¥2/M Qwen3.8-Max: ¥24/M, Qwen3.5-Max: ¥12/M, Qwen3.5-Plus: ¥6/M ✅ 国内直连 新用户 100万 tokens 免费额度(90天有效);Qwen3.5 系列开源可自部署 查看详情 →
Perplexity AI Sonar Pro, Sonar Sonar Pro: $3/M, Sonar: $1/M Sonar Pro: $15/M, Sonar: $1/M ❌ 需代理 API 无免费额度,Pro 订阅 $20/月(5000次搜索查询) 查看详情 →
美团 LongCat LongCat-2.0 (1.6T MoE), LongCat-2.0 INT4 OpenRouter: $0.60/M token (48B active MoE) OpenRouter: $2.40/M token (48B active MoE) ✅ Domestic self-host + OpenRouter (proxy) MIT licensed weights on Hugging Face; self-host on 8x H100+ GPU cluster 查看详情 →
Meta Model API Muse Spark 1.1 (muse-spark-1.1) $1.25/1M tokens $4.25/1M tokens ❌ 需代理(Meta Model API 目前仅面向美国开发者公开预览,中国大陆需代理访问) Free tier: 60 RPM, 2M TPM limit; free credits on signup for US developers (public preview) 查看详情 →
阶跃星辰(Stepfun) Step-2 (1.2T multimodal flagship, text+image+audio+video, 128K context), Step-2-mini (200B multimodal, faster Step-2 variant) Step-2: ¥6/M, Step-2-mini: ¥1/M, Step-1: ¥4/M, Step-R: ¥8/M Step-2: ¥18/M, Step-2-mini: ¥3/M, Step-1: ¥12/M, Step-R: ¥24/M ✅ 国内直连(platform.stepfun.com 端点,国内BGP优化) 新用户 100万 tokens 免费额度(注册后30天有效) 查看详情 →
Exa Search — neural semantic search across the web; 6 search types (auto / instant / fast / deep-lite / deep / deep-reasoning); returns 10 results by default, $1/1k per extra result above 10; AI page summaries $1/1k pages, Contents — full page text / highlights / summaries for known URLs; $1/1k pages per content type (text, highlights, summary billed separately) Pure pay-as-you-go by endpoint and search-type. Search $7/1k requests (base 10 results) + $1/1k per extra result + $1/1k AI summaries; Contents $1/1k pages; Answer $5/1k; Monitors $15/1k; Deep Search $12-15/1k; Agent fixed effort $0.012-$1.00/request or usage-based $0.10/ACU + tool calls (default $5 auto cap, $20 max cap). New accounts get $20 in free credits (~2,800 searches); Free Tier adds $10/month. No subscription, no minimum spend. Same as input. Enterprise custom volume + Zero Data Retention + SLA + postpaid invoice. ❌ Proxy required. Exa primary API runs at https://api.exa.ai from a single-region US deployment (San Francisco). No documented mainland China endpoint as of 2026-08-30. Mainland China production traffic typically needs a proxy or relay; transpacific first-byte latency is usually 100-300 ms. Enterprise plans can negotiate custom routing and dedicated infrastructure for regulated workloads (HIPAA / GDPR data residency). dashboard.exa.ai/onboarding provides agent-friendly onboarding that generates integration snippets tailored to the developer's stack, but API key issuance still requires stable international network access from China. Permanent $20 signup credit (~2,800 Search calls) + automatic $10 Free Tier monthly credit refill. No credit card, no minimum spend. Free credits apply to every endpoint. 查看详情 →
阿里云百炼 Qwen3.5-Max, Qwen3.5-Plus Qwen3.5-Max: ¥4/M, Qwen3.5-Plus: ¥2/M, Qwen3.5-72B: ¥4/M, QwQ: ¥2/M Qwen3.5-Max: ¥12/M, Qwen3.5-Plus: ¥6/M, Qwen3.5-72B: ¥12/M, QwQ: ¥8/M ✅ 国内直连 新用户 100万 tokens 免费额度(90天有效) 查看详情 →
百度文心一言 ERNIE 4.5 Turbo, ERNIE 4.0 Turbo ERNIE 4.5 Turbo: $0.003/M tokens, ERNIE 4.0 Turbo: $0.012/M ERNIE 4.5 Turbo: $0.003/M, ERNIE 4.0 Turbo: $0.012/M ✅ 国内直连 ERNIE Speed/Lite/Tiny 免费调用,每个模型每月 10 万 tokens 查看详情 →
Vercel AI Gateway OpenAI GPT-4o / GPT-4o-mini / GPT-5.x, Anthropic Claude Opus 4.8 / Sonnet 4 免费层: $5/月 AI Gateway Credits;付费层: 按量付费,Provider 官价零加价 Token 零加价(含 BYOK);附加功能另计 ⚠️ 需代理(境内需 VPN) 每团队每月 $5 免费 AI Gateway Credits(仅限免费层可用模型) 查看详情 →
月之暗面 Kimi Kimi K3, Kimi K2.7 Code K3: ¥2/M cached, ¥20/M uncached; K2.7/K2.6: ¥1.10-¥1.30/M cached, ¥6.50/M uncached K3: ¥100/M; K2.7/K2.6: ¥27/M ✅ 中国大陆直连 新用户 ¥15 代金券不可用于 Kimi K3;需充值后体验 查看详情 →
Mistral AI Mistral Large 2, Mistral Small Large 2: $2/M, Codestral: $1/M, Ministral 8B: $0.10/M Large 2: $6/M, Codestral: $3/M, Ministral 8B: $0.10/M ❌ 需代理 注册送免费 API 额度(La Plateforme 每日速率限制内免费) 查看详情 →
Dify GPT-4o / GPT-4o-mini / o1 / o3 (via OpenAI), Claude 3.5 Sonnet / Haiku / Opus (via Anthropic) Sandbox 免费:200 GPT-4 调用 + 5MB vector 存储 + 10 文档 + 5,000 API 调用/月;Professional $59/workplace/月:100MB向量 + 100文档 + 50 工作流 Team $159/workplace/月:500MB向量 + 500文档 + 200 工作流 + 500K 触发事件;Enterprise 合同制(SSO/私有部署/SLA) ✅ 国内母公司 LangGenius,原生中文支持、阿里云一键部署、SOC 2/GDPR/ISO 27001 认证,国内云完全可用 Sandbox 免费层:200 GPT-4 调用 + 5MB vector 存储 + 10 文档 + 2 工作流 + 5,000 API 调用/月(30 天日志保留) 查看详情 →
智谱 AI GLM GLM-5.3, GLM-5.3-Flash GLM-5.3 (Workers AI): $1.40/M + $0.26/M cached; GLM-5.2: ¥8/M; GLM-5.3-Flash (Workers AI): $0.15/M; GLM-5: ¥4-6/M; GLM-4.7: ¥2-4/M; GLM-4.5-Air: ¥0.8/M; GLM-4.7-Flash: ¥0/M (free) GLM-5.3 (Workers AI): $4.40/M; GLM-5.2: ¥28/M; GLM-5.3-Flash (Workers AI): $0.50/M; GLM-5: ¥18-22/M; GLM-4.7: ¥8-16/M; GLM-4.5-Air: ¥2-8/M; GLM-4.7-Flash: ¥0/M (free) ✅ 国内直连 GLM-4-Flash 完全免费;注册送 500万 tokens 体验包 查看详情 →
Cohere Command R+, Command R7 Command R7: $2.50/M, Command R: $0.50/M, Embed: $0.10/M Command R7: $10/M, Command R: $1.50/M ❌ 需代理 免费 Trial API Key(2500次调用/月,限 Command R) 查看详情 →
Cloudflare AI Gateway 通过 Gateway 路由到 20+ Provider 任意模型 免费层: 100万次请求/月;超出 $0.20/100万次 + Provider 自身费用 按请求计费,不按 token 计;Gateway 本身加价 $0.20/100万请求 ✅ 全球 CDN(国内边缘节点通过合作伙伴运营) 每月 100 万次请求免费(Gateway 服务费) 查看详情 →
腾讯混元 Hunyuan Hy4 preview, Hunyuan Turbo Turbo S: ¥0.8/M tokens, Turbo: ¥1.2/M, Lite: ¥0.3/M Turbo S: ¥0.8/M tokens, Turbo: ¥4.8/M, Lite: ¥0.3/M ✅ 国内直连 新用户 100万 tokens 免费额度(180天有效) 查看详情 →
Fireworks AI Llama 3.3 70B, Firefunction-v2 Firefunction-v2: $0.90/M, Llama 3.3 70B: $0.90/M, DeepSeek-R1: $2.00/M Firefunction-v2: $0.90/M, Llama 3.3 70B: $0.90/M, DeepSeek-R1: $2.00/M ❌ 需代理 新用户 $0.50 免费额度 查看详情 →
Portkey OpenAI GPT-4o / GPT-5 / GPT-5.6 family, Anthropic Claude Opus 5 / Sonnet 5 / Haiku Free 层: 10,000 请求/月;Hobby $49/月(10万请求);Growth $249/月(100万请求);Enterprise 合同制 请求量计费,不收 token 路由费;BYOK 免费;日志/可观测性/Caching 按订阅档位开放 ❌ 需代理(portkey.ai 域名在国内访问不稳定,推荐自托管开源版本) 每月 10,000 次请求;100K 日志保留;社区 Slack 支持;BYOK 免费;无信用卡要求 查看详情 →
Black Forest Labs FLUX.2 [pro], FLUX.1.1 [pro] FLUX.2 [pro]: $0.05/MP, FLUX.1.1 [pro]: $0.04/MP, FLUX.1 [schnell]: $0.003/MP, FLUX.1 Fill: $0.05/MP Flat per-megapixel billing, no tier-based markup ❌ 需代理(BFL API 部署在 AWS US-East + EU-Frankfurt,国内访问需稳定代理) Free tier via API: limited generations per minute for [schnell] and [dev] endpoints; [pro] requires paid credits 查看详情 →
字节豆包 Doubao Doubao-Seed-2.0 (旗舰推理), Doubao-Seed-2.0-lite (高性价比) Seed-2.0: ¥0.8/M, Seed-2.0-lite: ¥0.5/M, Seed-2.0-mini: ¥0.15/M Seed-2.0: ¥2/M, Seed-2.0-lite: ¥1/M, Seed-2.0-mini: ¥0.6/M ✅ 国内直连 新用户 50万 tokens 体验包 查看详情 →
LiteLLM 通过 Proxy 转发 100+ Provider 任意模型 开源免费自部署;LiteLLM Cloud 按量计费 开源免费自部署;LiteLLM Cloud 按量计费 ✅ 开源自部署,无限制 开源版完全免费(自托管) 查看详情 →
硅基流动 Qwen/Qwen3.5-Plus, Qwen/Qwen2.5-72B-Instruct ¥0.4-2/1M tokens (Qwen2.5-7B ~¥0.4,Llama 3.3 70B ~¥2) ¥0.4-2/1M tokens ✅ 国内直连 注册送 ¥1-200 免费额度(按活动调整),足够跑 1-20M tokens 查看详情 →
Jina AI jina-embeddings-v3, jina-embeddings-v2-base-en Embeddings v3: $0.02/M tokens (10K free); Reranker v3: $0.018/M tokens (10K free); Reader: $0.02/M tokens (free up to 1M tokens); CLIP-v1: $0.02/M tokens Flat per-million-token billing, no tier-based markup; batch discount of 50% available on all endpoints ❌ 需代理(Jina API 部署在 AWS US-East + EU-Frankfurt,国内访问需稳定代理) 10,000 free tokens/month on Embeddings + Reranker; Reader API free up to 1M tokens/month; no credit card required to start 查看详情 →
Arize Phoenix OpenTelemetry 原生 trace collector,支持任意 LLM 框架, Phoenix Cloud 托管 + 自托管 (Apache-2 / Elastic-2.0) Cloud 免费层: 10 GiB 存储/工作空间;自托管: $0(仅基础设施费) Phoenix Cloud 付费: 按量付费(存储计费);Arize AX 企业: 合同制 ✅ 开源自部署 + Cloud 托管(境外服务,境内需代理访问 SaaS) Phoenix Cloud 免费套餐:每工作空间 10 GiB 存储(无时间限制) 查看详情 →
Block AI(b.ai) GPT-4o, GPT-4o-mini 主流模型 + 0-10% 加价(如 GPT-4o $2.5→$2.75/M) 主流模型 + 0-10% 加价 ✅ 国内直连 注册送体验额度(每天免费请求次数待确认) 查看详情 →
Groq Llama 3.3 70B, Llama 3.1 405B Llama 3.3 70B: $0.59/M, DeepSeek-R1: $0.89/M, Mixtral 8x7B: $0.15/M Llama 3.3 70B: $0.79/M, DeepSeek-R1: $0.89/M, Mixtral 8x7B: $0.15/M ❌ 需代理 免费速率限制(足够个人开发和小型测试使用) 查看详情 →
fal.ai FLUX.2 [pro] / [dev] / [schnell], Kling Video 2.1 / 2.0 FLUX.2 [pro]: $0.05/MP (image), $0.08/sec (video); Kling 2.1: $0.10/sec; HunyuanVideo 1.5: $0.08/sec Flat per-second / per-megapixel billing, no tier-based markup ❌ 需代理(fal.ai 部署在 AWS US-East,国内访问需稳定代理) $1 free credit on signup (no credit card required); 100+ models available free for first-time evaluation 查看详情 →
Helicone 通过 AI Gateway 代理 100+ Provider 任意模型, AI Gateway + LLM Observability 集成 免费层: 10,000 请求/月 + 1GB 存储;Pro $79/月(团队版含告警、HQL 查询);Team $799/月(SOC-2/HIPAA) 按月订阅 + 用量计费(超过免费额度) ✅ 开源自部署 + SaaS 每月 10,000 次免费请求 + 1GB 存储 (Hobby) 查看详情 →
FreeModel DeepSeek V3 / R1, Qwen 2.5 / QwQ-32B 模型相关(平台加价待确认) 模型相关 ✅ 国内可用(直连) 注册送额度(具体金额待确认) 查看详情 →
Cerebras Cerebras Llama 3.3 70B, Cerebras Llama 3.1 405B $0.10/M to $0.60/M tokens(所有模型统一 $0.60/M input) $0.10/M to $0.60/M tokens(统一 $0.60/M output) ❌ 需代理 免费速率限制(限制 RPM,适合测试) 查看详情 →
AI21 Labs (Jamba) Jamba 1.5 Mini, Jamba 1.5 Large Jamba 1.5 Mini: $0.20/M tokens, Jamba 1.5 Large: $2/M tokens Jamba 1.5 Mini: $0.20/M tokens, Jamba 1.5 Large: $2/M tokens ❌ 需代理 无免费额度 查看详情 →
CoreWeave NVIDIA HGX H100 ($49.24/hr), HGX H200 ($50.44/hr), HGX B200 ($68.80/hr), A100 80GB ($21.60/hr), L40S ($18.00/hr), L40 ($10.00/hr), GB200 NVL72 ($42.00/hr) — On-Demand/Spot GPU instances, Serverless Inference (W&B Inference, pay-per-token): GLM 5.2, Kimi K2.6/K2.7, DeepSeek R1/V3, Llama 3.x, Qwen — OpenAI-compatible API GPU On-Demand hourly: NVIDIA HGX H100 $49.24, HGX H200 $50.44, HGX B200 $68.80, A100 80GB $21.60, L40S $18.00, L40 $10.00, GB200 NVL72 $42.00. Spot discounts: H100 $19.71, H200 $20.93, B200 $34.11, A100 $9.65. Serverless Inference is pay-per-token via W&B Inference; Dedicated Inference is GPU-hour / node-based. ❌ 无中国大陆直连节点。数据中心集中在美国;2026 年已扩展至印尼(首个 APAC 节点)、英国(两个数据中心运营中)、瑞典。国内访问需代理,跨太平洋延迟较高。 无永久免费层。Serverless Inference 按 token 计费;GPU 实例按小时计费(On-Demand)或有 spot 折扣。0EM(零出口费迁移)计划在迁移阶段免除出口费。 查看详情 →
APIKEY.FUN GPT-4o, Claude 3.5 Sonnet 主流模型,价格约为官方 API 的 60-80% 主流模型,价格约为官方 API 的 60-80% ✅ 国内直连 注册送少量免费额度 查看详情 →
DeepInfra Meta-Llama-3.3-70B-Instruct, Meta-Llama-3.1-405B-Instruct Llama 3.3 70B: $0.49/M, DeepSeek-V3: $0.99/M, DeepSeek-R1: $1.09/M Llama 3.3 70B: $0.73/M, DeepSeek-V3: $0.99/M, DeepSeek-R1: $1.09/M ❌ 需代理 注册即送免费额度(每日有限免费调用) 查看详情 →
NVIDIA NIM Nemotron-3 Ultra 550B (a55B), Nemotron-3 Super 120B (a12B) Partner-dependent (Nemotron-3 Ultra 550B: $0.50-0.90/1M tokens) Partner-dependent (Nemotron-3 Ultra 550B: $1.70-3.60/1M tokens) ❌ 需代理(build.nvidia.com 平台为美国服务,受出口管制限制) 免费开发端点(无需信用卡,生成 API Key 即可开始) 查看详情 →
Amazon Bedrock Claude Opus 4.8 / 4.7 / 4.6 / 4.5, Claude Sonnet 4.6 / 4.5 / 4 Claude Sonnet 4.5: $6/M, DeepSeek V3.2: $0.62/M, Mistral Large 3: $0.50/M, GPT-5.5: $5.50/M, GPT-5.4: $2.75/M, Grok 4.3: $1.25/M, Qwen3 235B: $0.23/M, Kimi K2.5: $0.60/M Claude Sonnet 4.5: $30/M, DeepSeek V3.2: $1.85/M, Mistral Large 3: $1.50/M, GPT-5.5: $33/M, GPT-5.4: $16.50/M, Grok 4.3: $2.50/M, Qwen3 235B: $0.91/M, Kimi K2.5: $3/M ❌ 大陆用户需代理(AWS 中国区(由光环新网/西云数据运营)尚未提供 Bedrock 服务,需使用全球区域 + VPN/代理) AWS Free Tier 赠送新用户最高 $200 抵扣金(6个月有效)。Bedrock 按用量抵扣,无永久免费额度 查看详情 →
Hyperbolic DeepSeek R1, DeepSeek R1-0528 Per-token inference scaled from market GPU rates; H100 from $2.89/GPU-hour, H200 $3.49/GPU-hour Pay-as-you-go GPU compute billed per hour; reserved discounts for committed capacity ❌ 需代理 无长期合同;按量计费按小时结算;无永久免费档,但无配额上限、即开即用 查看详情 →
Fireworks AI DeepSeek V4 Pro — $1.74 input / $0.145 cached / $3.48 output per 1M tokens (Serverless Standard; Priority $2.61 / $0.218 / $5.22), DeepSeek V4 Flash (0731) — $0.22 input / $0.007 cached / $0.66 output per 1M (Standard) Serverless (Standard, $/1M tokens): DeepSeek V4 Pro $1.74、V4 Flash $0.22、Kimi K3 $3.00、K2.7 Code $0.95、MiniMax M3 $0.30、Qwen 3.7 Plus $0.40、GPT OSS 120B $0.15、GLM 5.1 $1.40。嵌入按模型档 $0.008-$0.016/1M。按需 GPU 按小时:H100 $7.00、H200 $7.00、B200 $10.00、B300 $12.00、GB300 $18.00(9 月 1 日上调)。 Serverless 输出($/1M):DeepSeek V4 Pro $3.48、V4 Flash $0.66、Kimi K3 $15.00、MiniMax M3 $1.20、Qwen 3.7 Plus $1.60、GPT OSS 120B $0.60、GLM 5.1 $4.40。缓存输入显著更低(如 DeepSeek V4 Pro $0.145)。训练:SFT 每 1M 训练 token $0.50(≤16B)到 $10.00(>300B);Serverless Training API 仅按 token 计费、无闲置 GPU 成本。 ❌ 无中国大陆直连端点。作为美国平台,大陆直连需代理或海外中转,跨太平洋首字节延迟通常 150-250ms。最近 APAC 端点为东京多区域。 注册即享 $1 免费 serverless 额度(pay-per-token,postpaid billing);无永久免费档。按需 GPU 与训练按用量计费。 查看详情 →
Replicate Llama 3.3 70B, DeepSeek-R1 模型调用按运行时长 + GPU 型号计费,$0.00025/s/A100起步 按运行时长计费(非 token 计费模式) ❌ 需代理 新用户 $0.50 免费额度 查看详情 →
SambaNova Llama 3.3 70B, Llama 3.1 405B Llama 3.3 70B: $0.50/M, DeepSeek-R1: $0.80/M, Qwen 72B: $0.40/M Llama 3.3 70B: $0.50/M, DeepSeek-R1: $0.80/M, Qwen 72B: $0.40/M ❌ 需代理 免费速率限制(测试用途) 查看详情 →
零一万物 Yi Yi-Lightning (智能路由 → DeepSeek-V3/Qwen3-30B-A3B/Yi-Lightning), Yi-Vision-v2 (视觉理解,路由 → Qwen2.5-VL-72B/Yi-Vision-V2) ¥0.99/1M total tokens(Yi-Lightning,input+output 合并计费) ¥0.99/1M total tokens(合并计费,不区分 input/output) ✅ 国内直连(国内公司,支持手机号注册 + 支付宝/微信支付) 注册即开通免费速率限制(有限 RPM/TPM),无初始赠送余额 查看详情 →
Nebius DeepSeek V4 Flash, DeepSeek V4 Pro Per-token OpenAI-compatible inference: DeepSeek-V4-Flash $0.14/M input, $0.28/M output; Kimi K3 $3/$15; MiniMax M3 $0.30/$1.20; Llama 3.3 70B $0.13/$0.40 Two flavors per model - Base and Fast (-fast suffix) - same outputs, Fast trades higher token price for lower latency via speculative decoding ❌ 需代理 注册即可起步,无月费;按量付费(pay-as-you-go)按 token 计费;弹性动态限流,随用量自动扩容至基础额度 20 倍,无需预购容量 查看详情 →
Anyscale Llama 3.3 70B, Llama 3.1 405B Llama 3.3 70B: $0.39/M, DeepSeek-R1: $1/M, Mixtral 8x22B: $0.90/M Llama 3.3 70B: $0.39/M, DeepSeek-R1: $1/M, Mixtral 8x22B: $0.90/M ❌ 需代理 新用户 $100 免费额度(一次性),另提供少量项目启动额度($3-$5) 查看详情 →
DigitalOcean Gradient Llama 3.3 70B, Llama 3.1 405B Llama 3.3 70B: $0.50/M, DeepSeek-R1: $0.80/M Llama 3.3 70B: $0.50/M, DeepSeek-R1: $0.80/M ❌ 需代理 新用户 $100 免费额度(60 天有效) 查看详情 →
MiniMax(海螺 AI) MiniMax-01 (Lightning / Turbo / Pro), MiniMax-Text-01 MiniMax-01 Lightning: ¥0.2/M tokens, Turbo: ¥0.8/M, Pro: ¥2/M, M3: ¥1.5/M MiniMax-01 Lightning: ¥0.2/M tokens, Turbo: ¥0.8/M, Pro: ¥2/M, M3: ¥1.5/M ✅ 国内直连 注册即送 100 万 tokens 免费额度(90 天有效期) 查看详情 →
Stability AI Stable Diffusion 3.5, Stable Diffusion XL SD3.5: $0.0065/image, SDXL: $0.01/image, Stable Code: $0.50/M tokens 按图像/视频/音频输出计费,非 token 计费 ❌ 需代理 新用户 25 次免费生成 查看详情 →
Novita AI Llama 3.3 70B, Llama 3.1 405B Llama 3.3 70B: $0.59/M, DeepSeek-V3: $1.00/M Llama 3.3 70B: $0.79/M, DeepSeek-V3: $1.00/M ⚠️ 部分可用(新加坡节点,国内延迟尚可) $0.50 免费试用额度 查看详情 →
Baseten Llama 3.3 70B, Llama 3.1 405B Per-second GPU billing (H100: $1.89/hr, A100: $0.83/hr, L40S: $0.54/hr) Model-dependent (inference time × GPU rate) ⚠️ 部分可用(需稳定代理连接) $30 in inference credits on signup (valid 30 days); custom model hosting free tier: 1 deployed model 查看详情 →
Mem0 OpenAI GPT-4o / GPT-4-Turbo / GPT-3.5, Anthropic Claude 3.5 Sonnet / Haiku / Opus Hobby 免费:10,000 memories + 1,000 retrieval API calls/月;Starter $19/月:50,000 memories + 5,000 calls Pro $249/月:Unlimited memories + 50,000 retrieval API calls + Graph Memory + 多项目;Enterprise 合同制 ✅ 开源(Apache-2.0,61k+ stars)+ Cloud SaaS;国内自部署无障碍 Hobby 免费层:10,000 memories + 1,000 retrieval API calls/月(无限 end users),社区支持 查看详情 →
Tavily Search (real-time web + AI summary, 1 credit/call), Extract (clean markdown from URL, 1 credit/page) Free 1,000 credits/month(滚动 30 天窗口,无信用卡);Pay-as-you-go $0.008/credit(基础 Search 1 credit);Project 起步 $30/month 含 4,000 credits,$0.008/credit 超额;Growth $200/month 含 40,000 credits,$0.006/credit 超额;Pro $1,000/month 含 250,000 credits,$0.005/credit 超额;Enterprise 合同制(custom 配额 + SLA + 无日志合约) 按 endpoint 计费:Search 1 credit/次;Extract 1 credit/页;Crawl 1 credit/页;Map 1 credit/页;Research 5 credits/次(替代 5-10 Search + 2-3 Extract + 1 合成 LLM);AI Extraction 1 credit/次 ⚠️ Tavily API 部署在 AWS US-East / EU-West;国内直连延迟 200-400ms;推荐用 Cloudflare Worker 或腾讯云 Edge Function 中转(降至 50-100ms);无 CN 节点计划;无国内信用卡支付渠道 Free 永久:1,000 credits/month(滚动 30 天窗口),无信用卡、无功能限制;Search/Extract/Crawl/Map/AI Extraction 各占独立配额(1,000 Search + 500 Crawl);超额返 HTTP 429,无超额计费 查看详情 →
Lepton AI Llama 3.3 70B, Llama 3.1 405B Llama 3.3 70B: $0.80/M; Llama 3.1 405B: $3.50/M; Qwen 2.5 72B: $0.80/M; DeepSeek-R1: $2.00/M Same rate as input for most models (symmetric pricing) ⚠️ 部分可用(ap-northeast-1 东京 region 距离最近,国内访问需代理) $5 试用额度(30 天有效,足够完整评估任意一款主机模型) 查看详情 →
可灵 AI(快手) Kling 3.0 (native 4K video, multi-shot sequencing, Native Audio), Kling 3.0 Omni (multimodal video, video/image reference input) Per-second video billing. Kling 3.0 Turbo: 720p $0.112/s, 1080p $0.14/s. Kling 3.0: 720p $0.084-0.126/s, 1080p $0.112-0.168/s (Native Audio), 4K $0.42/s. Kling 3.0 Omni: 720p $0.084-0.126/s, 1080p $0.112-0.168/s, 4K $0.42/s, varies by video/audio input. Image API: Kling Image 3.0 $0.028/image (1K/2K), 3.0-omni $0.028-0.056/image (up to 4K), Image 2.1 $0.014-0.028/image, Image O1 $0.028/image, Multi-Shot $0.07/call. Video list price 1 Unit = $0.14; image list price 1 Unit = $0.0035. ✅ 官方支持部分中国区域调用(快手自研,国内云节点可用;海外 klingai.com 需网络环境) 新用户注册后可获得少量免费体验积分(平台内赠送 credits),可体验部分视频/图片生成;正式 API 使用无永久免费套餐,按生成的秒数/张数计费。 查看详情 →
Hugging Face Llama 3.3 70B, Llama 3.1 405B 推理 API: $0.04-0.60/M tokens 取决于模型 推理 API: $0.04-0.60/M tokens ❌ 需代理(huggingface.co 在国内被墙) 免费推理 API(有限速率限制) 查看详情 →
ElevenLabs Eleven Multilingual v2, Eleven Turbo v2.5 TTS: $0.03-0.30/1K chars depending on model/quality; STT: $0.47/hour; Sound Effects: $0.04/request Same as input rate per model ⚠️ 部分可用(需稳定代理) 10,000 credits/month (≈10 minutes audio, free plan with limited features) 查看详情 →
Ideogram Ideogram 3.0, Ideogram Turbo Ideogram 3.0: $0.04/image (standard), $0.08/image (Turbo) 4 images per generation (default), additional images via paid credits ❌ 需代理 新用户赠送 $1 体验金,约 25 次标准生成 查看详情 →
Lambda GPU Cloud NVIDIA HGX B200 180GB (on-demand GPU instances; $6.69/GPU/hr), NVIDIA H100 SXM 80GB ($3.99/GPU/hr on-demand) GPU instance on-demand per GPU/hr: NVIDIA B200 SXM6 $6.69, H100 SXM $3.99, A100 SXM 80GB $2.79, A100 SXM 40GB $1.99, GH200 $2.29, A100 PCIe $1.99, A6000 $1.09, A10 $1.29, V100 $0.79. 1-Click Cluster reserved per GPU/hr: B200 $9.86 (16 GPU) to $8.87 (256+), H100 $6.16 to $5.54. Superclusters / Private Cloud (4000+ GPUs) via sales. No egress fees. Billed by the second per GPU instance (no idle GPU cost); clusters and reserved capacity carry a minimum commitment (2 weeks to 1 year). Managed orchestration: Managed Kubernetes and Slurm included at no extra per-node fee; Lambda Stack one-line install. No per-token inference pricing — Lambda's Inference API is winding down; self-hosted inference billed as GPU runtime. ❌ 无中国大陆直连节点;数据中心在美国(加州等区域),国内访问需代理或海外中转。 无永久免费 GPU 层;注册后按用量计费(GPU 实例按秒计费)。1-Click Clusters 与 Superclusters 需预付 / 承诺周期(2 周至 1 年)。 查看详情 →
Deepgram Nova-3 (STT, Monolingual), Nova-3 (STT, Multilingual) STT: $0.0043-0.012/min; TTS: $0.015-0.030/1K chars N/A ⚠️ 部分可用(需代理,无中国大陆数据中心) New users get $200 in free credits; no credit card required 查看详情 →
RunPod B300 288GB HBM3e, B200 180GB Per-second GPU: H100 SXM $2.99/hr, A100 SXM $1.49/hr, B200 $5.89/hr, L40S $0.99/hr, RTX 4090 $0.69/hr Same as input (per-GPU-second — no idle fees, no egress fees) ⚠️ 部分可用(基础设施在全球 31 区域,国内直连不稳定,需稳定代理) $5 starter credit on signup, no credit card required; per-second billing means spiky workloads never pay for idle 查看详情 →
Luma AI Ray 3.2 (flagship video, 1080p, multi-keyframe up to 16 frames, V2V up to 20s), Ray 3.14 (previous-gen video model) Credit-based subscription. Plans: Plus $30/mo (10,000 credits), Pro $90/mo (40,000 credits), Ultra $300/mo (150,000 credits). Ray 3.2 video: 1080p text-to-video 400 credits/5s, 720p 100 credits/5s, Draft 20 credits/5s; Seedance 2.0 1080p 240 credits/sec, 4K 959 credits/sec. Image credits: Uni-1 30 credits/image, Seedream 1-3 credits/image, GPT Image 2 from 3 (Low-1K) to 255 (High-4K) credits. Audio: ElevenLabs v3 TTS 21 credits/1,000 chars. Utilities: background removal 1 credit/image, reframe video 32 credits/sec, upscale to 4K 17 credits/sec. ❌ 需代理(无中国大陆直连节点,API 托管于全球 AWS/GCP 边缘) 无免费 API 套餐。Luma 提供有限免费信用额度供新用户体验(平台注册后少量免费 credits),但正式 API 使用需至少订阅 Plus $30/月(10,000 credits)。 查看详情 →
Chroma All-MiniLM-L6-v2 (default, 384-dim, SBERT), All-mpnet-base-v2 (768-dim, higher quality SBERT) Chroma OSS 完全免费(Apache 2.0),本地/嵌入式运行;Chroma Cloud Free 永久层:50K 向量、5K 查询/月、0.5 GB 存储;Pro $0.30/M 向量-月 + $0.01/M 查询(月最低 $10);Enterprise 合同制(自定义节点、HIPAA、SOC 2) 存储(GB-月)+ 向量数 + 查询数 三维度计费;无 per-embedding API 费;Default Embedding Function 本地推理零成本;BYO Embedding 由第三方 API 计费 ✅ Chroma OSS 可完全本地运行(国内服务器、零延迟);Chroma Cloud 节点在 AWS us-east-1 / eu-west-1 / ap-southeast-2(2026 H2 新加坡节点待上线),直连大陆延迟 200-400ms Chroma OSS 永久免费、Apache 2.0、无功能限制、无向量数限制(本地硬件决定);Chroma Cloud Free 永久层:50K 向量 + 5K 查询/月 + 0.5 GB 存储 + 1 个 Collection + 社区支持 查看详情 →
Voyage AI voyage-4-large, voyage-4 voyage-4-large: $0.12/M, voyage-4: $0.06/M, voyage-4-lite: $0.02/M, voyage-context-3: $0.18/M, voyage-code-3: $0.18/M, voyage-multimodal-3.5: $0.12/M (embedding API — no output tokens; rerank-2.5: $0.05/M tokens, rerank-2.5-lite: $0.02/M tokens) ⚠️ 部分可用(官网在中国可访问,但 API 在中国大陆的稳定性未公开声明) 200M free tokens per account (text embedding + reranker), 50M free tokens for legacy domain models, 200M free text tokens + 150B free pixels for multimodal 查看详情 →
Runway Gen-4 Turbo, Gen-4 Alpha Gen-4 Turbo: 50 credits/5s clip (~$0.50), Gen-4 Alpha: 100 credits/5s (~$1.00), Act-Two: 150 credits/5s (~$1.50), Frames: 200 credits/sequence (~$2.00) Credit-based system (~$0.01/credit). Credit packs: $12 (1,000 credits) to $200 (25,000 credits). Enterprise: custom pricing. ❌ 需代理(AWS US-East + GCP US-Central 托管) 无免费 API 套餐。最低 $12 购买 1,000 credits,约 20 次 Gen-4 Turbo 生成。Web 应用免费但无 API 访问。 查看详情 →
TwelveLabs Pegasus 1.5 (flagship multimodal video understanding: frames + audio + speech + on-screen text), Pegasus 1.2 (previous-gen, lower-cost option: frames + audio + speech) Free plan: 600 minutes video indexing + API usage (Pegasus 1.2 $0.042/min indexing one-time, $0.021/min input video, $0.0075/1k tokens output text, $0.0015/min monthly embedding infrastructure). Developer plan: paid overage, 90-day index retention lifted to unlimited. Same per-minute and per-1k-tokens billing. Marengo embeddings billed separately by indexed minute. ❌ 需代理(无中国大陆直连节点,API 托管于 AWS us-west-2 / us-east-1 区域) 免费计划:注册即获 600 分钟视频索引额度(一次性累积额度),可调用全部核心 API(Pegasus 1.2/1.5、Marengo 3.0、Search、Embed、Analyze & Segment)。索引数据保留 90 天;升级 Developer 计划后索引保留变为无限。 查看详情 →
Mixedbread AI Toast 1 — specialized search model (2026-08-13): matches/outperforms Claude Opus 5 and GPT-5.6 Sol on knowledge work, up to 10x cheaper and 12x faster; input $0.50 / cached $0.06 / output $1.20 per 1M LLM tokens (launch $0.30 / $0.036 / $0.72), mxbai-embed-large-v1 — 1,024-dimension open embedding model (Apache-2.0), 50M+ downloads across the mxbai family 按量计费,三块:索引(Fast $1.50/百万内容 token;High Quality 含 OCR/转写/摘要/多模态增强 $3/百万)、搜索(Grep $0.10/千次、Semantic $4/千次、Toast 1 $1/千次,重排 +$3.50 或 +$1.50/千次)、存储 $0.50/百万内容 token/月。Toast 1 按 LLM token:输入 $0.50/百万(发布价 $0.30)。 Toast 1 输出 $1.20/百万 LLM token(发布价 $0.72,40% 折扣);缓存输入 $0.06/百万(发布价 $0.036),缓存写入免费。Semantic 检索重排加价 $3.50/千次。 ❌ 无中国大陆直连端点。Mixedbread 是美国/欧洲平台(法人实体 mixedbread ai inc.,柏林/旧金山),大陆直连需代理或海外中转,跨太平洋首字节延迟通常 150-250ms。 Starter 免费送 $5 一次性额度,无需信用卡;3 个工作区用户、10 个 Store、每分钟 100 次请求,适合评估真实语料。无永久免费档。 查看详情 →
Crusoe Cloud DeepSeek V3 0324 (legacy frontier: $0.50 input / $1.50 output per 1M tokens), DeepSeek V4 Flash (efficient frontier: $0.14 input / $0.28 output per 1M tokens, 70B tier) Pay-as-you-go per 1M tokens, four tiers by parameter count: <16B ($0.40) / 16B-70B ($2.50) / 70B-300B ($6.00) / >300B ($10.00). Serverless Fine-Tuning follows identical 4-tier structure. Managed Inference spans DeepSeek V3/V4 Pro/V4 Flash, GLM 5.1/5.2, GPT-OSS 20B/120B, Gemma 4 31B-it, Kimi K2.6, Llama 3.1/3.3, Nemotron 3 family (VoiceChat/Ultra/Lightning/Nano/Super/Nano Omni), Qwen3 family (8B/235B/Qwen3.5/Qwen3.6), Yutori n1.5. Cached tokens at $0.03-$1.50 per 1M. Self-Serve Deployments: NVIDIA H100 80GB HGX $5.50/hr, NVIDIA H200 141GB HGX $6.00/hr (dedicated endpoints for open + fine-tuned models). Tailored Deployments and Provisioned Throughput via sales engagement. Managed Kubernetes $0.10/cluster hour; Container Registry $0.10/GiB-month; Object Storage $0.06/GiB-month. ❌ 无中国大陆直连节点。托管推理 API 托管于 Crusoe 美国数据中心(Colorado / Texas 等);国内访问需代理。 无永久免费层;按使用量计费;可联系销售获取批量折扣。Serverless Fine-Tuning 四档起价 $0.40/M tokens。 查看详情 →
Letta Letta Agents SDK — open-source stateful agent runtime (Apache-2.0), with first-class persistent memory (core/archival/recall blocks) and runtime memory editing, Letta Code (`@letta-ai/letta-code`, Node.js 22.19+) — memory-first CLI coding agent, #1 on Terminal-Bench (Dec 2025 launch) 按用量:Free $0(最多 3 个有状态 agent、BYOK)、Pro $20/月(最多 20 个 agent、Letta Auto 周/月额度 + 按量超额)、Teams Pro 按席位(共享 agent、权限)、Developer Plan 纯 API 调用按信用(无 agent 上限)、Enterprise 商务报价(自托管/定制模型)。服务端工具按 CPU 时间 $0.00015/秒计费。 Letta Auto 用量超出 Pro 套餐限额后按 pay-as-you-go 模型价(参考 platform.letta.com/models)。Mods、Skills、Conversations 等功能均包含在套餐内,不单独计费。远程 MCP 工具由 MCP 供应商执行,不消耗 Letta 信用。 ❌ 无中国大陆直连端点。Letta, Inc. 注册于美国旧金山,platform.letta.com 是单一全球端点;大陆生产流量需要代理或中转,跨太平洋首字节延迟通常 150-250ms。Agent harness 完全开源(Apache-2.0),可自托管在自有基础设施上规避直连限制。 Free $0/月(最多 3 个有状态 agent),自带 API key;Letta Auto 用量与 Skills 受限,可用于评估和小型个人助手。 查看详情 →
Weaviate Snowflake Arctic-embed-m-v1.5, Snowflake Arctic-embed-m-v2.0 Free 永久层:100,000 对象 + 1 集合 + 7 天备份;Flex 月付 $45 起 + 从 $0.00465/1M 维度 + $0.12/GiB 存储;Premium 预付从 $400/月 Dedicated 合同制 from $400/月 起;Vector dimension、Storage、Backup 三维度独立计费;Embeddings 按 token 用量计费 ⚠️ 服务器在 AWS/GCP 海外节点,直连中国大陆延迟较高 Free 永久层:100,000 对象、1 集合、7 天备份;2,000 Embedding 请求/日;Query Agent 1,000 请求/月 查看详情 →
Writer Palmyra X5, Palmyra X4 Palmyra X5: $0.60/M, Palmyra X4: $2.50/M Palmyra X5: $6.00/M, Palmyra X4: $10.00/M ⚠️ 部分可用(官网在中国可直接访问,但 API 稳定性未公开声明) 14-day free trial, no credit card required (Starter plan); no standalone free API credits 查看详情 →
Reka Reka Core — top-tier multimodal model (text + image + audio + video); input $2.00 / output $6.00 per 1M tokens, image $0.02 / video $0.08 / audio $0.02 per minute, Reka Flash — cost-efficient model for everyday tasks (text + image + audio + video); input $0.80 / output $2.00 per 1M tokens, image $0.01 / video $0.06 / audio $0.015 per minute Reka Edge $0.10/M 输入;Reka Flash $0.80/M;Reka Core $2.00/M;Reka Flash Research 按请求计费($25-$60/千次)。按量付费,无最低承诺。 Reka Edge $0.10/M;Reka Flash $2.00/M;Reka Core $6.00/M。多模态(图像、视频、音频)单独按"每分钟"计费,Flash 模型视频 $0.06/分钟、Core $0.08/分钟;图像按调用计费(Flash $0.01、Core $0.02)。 ❌ 无中国大陆直连端点。Reka 注册于美国加州 Sunnyvale(530 Lawrence Expressway),api.reka.ai 是单一全球端点;大陆生产流量需代理或中转,跨太平洋首字节延迟通常 150-250ms。开源权重可自托管(NVIDIA GPU ≥24GB VRAM),规避直连限制。 通过 platform.reka.ai 注册即可获取 API key;Pay-as-you-go 用量计费,不强制最低承诺,可在充值额度内自由试用。无永久免费档、无月赠送额度。 查看详情 →
Qdrant FastEmbed: BAAI/bge-small-en-v1.5, FastEmbed: BAAI/bge-base-en-v1.5 Free 永久层:1 GB 存储 + 0.5M 向量,无限期;Cloud Standard 月付 $25 起(预付 $250/yr)按存储 $0.06/GB-月;Cloud Pro 月付 $80 起 按存储 $0.04/GB-月 + 高 IOPS;Dedicated 合同制 $2,500/月起 按存储维度(GB)与可选性能包(Power Tiers)计费;无 API 调用费、无 per-vector 嵌入费;FastEmbed 按 token 数本地推理;BYO Embedding 由第三方 API 计费 ⚠️ Cloud 运行在 AWS Frankfurt / N. Virginia / Sydney 等海外节点,直连中国大陆延迟较高;OSS 可在腾讯云/阿里云国内区域自托管 Free 永久层:1 GB 存储 + 0.5M 1015-dim 向量;2 CPU/0.5 GiB RAM 实例;无限 API 请求;Community Discord 支持 查看详情 →
fal.ai FLUX.2 Pro (text-to-image, $0.03/MP first MP + $0.015/extra MP), FLUX1.1 [pro] (ultra-fast text-to-image, per-megapixel) 按请求/按时长计费,无月费。图像生成:FLUX.2 Pro $0.03/第一个 MP + $0.015/额外 MP、FLUX1.1 [pro] 按 MP、Nano Banana 2 $0.08/图、FLUX.1 [schnell] 无定价显示。视频:Kling v3 $0.084-0.168/秒、Seedance 2.0 $0.3034/秒(720p)。TTS:通过 ElevenLabs/MiniMax 等第三方提供。余额模式:预付 $10+ 存入信用余额,用完为止。开源模型按秒计费,秒级颗粒度。 ✅ 大陆直连可用(fal.ai 使用全球 CDN,无需代理);延迟因模型推理节点而异,典型 200-500ms 无免费 API 层(预付 $10 起买信用余额);无月费、无订阅、无隐藏费用;Serverless 模式下零运行时自动缩零 查看详情 →
Aleph Alpha Pharia-1-LLM-7B-control, Pharia-1-LLM-7B-control-aligned Enterprise contract — no public per-token price (quote-based; PhariaAI on-prem deployment custom-priced) Same — enterprise quote, no public per-token list ❌ 不适用(德国主权 AI 厂商,受欧盟出口管制约束) No public free tier — self-serve signup creates an account, but pricing is enterprise contract only 查看详情 →
Cartesia Sonic (real-time TTS), Sonic 2 (multilingual, voice cloning) Sonic: $0.03/1K chars (Standard), $0.06/1K chars (Pro), $0.12/1K chars (Turbo); Sonic 2: $0.04-0.15/1K chars TTS character-based pricing (output = audio), per-second for streaming ⚠️ 部分可用(需稳定代理) 10,000 characters/month free tier (≈10 minutes), no credit card required 查看详情 →
Hume AI EVI 3 (Empathic Voice Interface 3), OCTAVE TTS (voice design + cloning) EVI 3: ~$0.096/min (pay-per-second); OCTAVE TTS: $0.048/1K chars; Expression Measurement: $0.0008/sec Same as input (most endpoints symmetric) Proxy required (US export controls) Free credits on signup (amount per account), enough to trial EVI 3 查看详情 →
Modal B300 / B200 / H200 / H100 / A100 / L40S / A10 / L4 / T4 (GPU), Self-hosted LLM inference (vLLM, SGLang, TensorRT-LLM) 按 GPU 秒计费:H100 $0.001097/s、A100 80GB $0.000694/s、T4 $0.000164/s 同输入(按 GPU 时间,无 idle 费用) ❌ 需代理(基础设施在 AWS/GCP 海外区域,国内直连不稳定) Starter 计划免费 + $30/月 计算额度(每月刷新) 查看详情 →
AssemblyAI Universal-3.5 Pro (Async, 18 languages, native code switching, best diarization), Universal-2 (Async, 99 languages, balanced accuracy) 按音频小时计费:Universal-3.5 Pro $0.21/hr、Universal-2 $0.15/hr(pre-recorded) N/A(speech-to-text 计费基于音频时长,无 token 概念) ❌ 需代理(基础设施在 US,国内直连不稳定) 免费账户每月 185 小时 pre-recorded + 333 小时 streaming,无需信用卡 查看详情 →
Pinecone Serverless (Standard / Enterprise), Pod-based (s1 / p1 / p2) Serverless Standard: $50/月最低,$4-$4.50/M读单元,$16-$18/M写单元,$0.0005/ingestion单元 Serverless Enterprise: $500/月最低,$6-$6.75/M读,$24-$27/M写,$0.001/ingestion单元(multi-modal);Pod-based按节点小时计费 ⚠️ 国内访问 pinecone.io 需代理(SaaS-only,无开源版) 免费 Starter 计划(1个项目,Serverless索引,$50信用额度/月,有效期1年) 查看详情 →
Together AI DeepSeek V4 Pro, Qwen3.6-Plus / Qwen3.7-Max $0.03-$1.74/MTok(模型相关) $0.12-$4.50/MTok(模型相关) ⚠️ 需代理(美国服务商) 注册赠送 $1 初始额度(需绑定支付方式) 查看详情 →
FriendliAI GLM-5.2, GLM-5.1 GLM-5.2/5.1: $1.4/M, DeepSeek-V3.2: $0.5/M, Qwen3-235B: $0.2/M, MiniMax-M2.5: $0.3/M, Gemma-4-31B: $0.14/M GLM-5.2/5.1: $4.4/M, DeepSeek-V3.2: $1.5/M, Qwen3-235B: $0.8/M, MiniMax-M2.5: $1.2/M, Gemma-4-31B: $0.4/M ⚠️ Partially available — servers in US & Korea, ~200-400ms latency from CN No free tier — per-token billing with no minimum 查看详情 →
Suno v5.5 (latest model, royalty-free music), v5 (premium model) Free: $0 (10 songs/day, non-commercial); Pro: $10/mo (500 songs, commercial rights); Premier: $30/mo (2,000 songs, Studio) Annual: Pro $96/yr (20% off), Premier $288/yr (20% off); credit top-ups available as add-ons ⚠️ 国内访问受限(需稳定海外代理),无官方中国大陆节点 免费版每天 10 首歌曲(不可商用),新账户直接开通,无需信用卡 查看详情 →
Liquid AI LFM2.5-8B-A1B (MoE, 8B total / 1.5B active, 128K context) — open weights, free download / run / fine-tune under the LFM Open License (royalty-free until $10M annual revenue), LFM2.5-2.6B (dense, 128K context, agentic tool calling) — DSpark draft variant up to 3.18× GPU / 2.87× on-device decode speedup 无按 token 云 API 定价——LFM 权重免费下载/运行/微调(LFM Open License,公司年收入超过 1000 万美元前免版税)。推理要么在自有硬件上自托管,要么通过 LEAP SDK 本地部署;无 per-token 请求费用。 无 per-token 输出计费。模型在 CPU/GPU/NPU 上本地或自托管运行,适合对延迟与隐私敏感的高频或离线工作负载(Liquid 的定位:「为何为云 API 按 token 付费,而非本地推理」)。 ⚠️ 无托管云 API,因此无中国直连端点问题。模型权重自由下载并可在本地/自托管运行,包括在中国境内(对数据主权友好)。不需要云 API 网关或代理;token 处理完全在本地硬件上进行。 全部 14+ 个 LFM 模型在 LFM Open License 下免费下载/运行/微调——包括商用(公司年收入 ≤ 1000 万美元)。研究、教育与非营利使用永久免费、无收入上限。 查看详情 →
Nscale Moonshot AI Kimi K2.5 — open chat model via serverless inference; per-1M input/output token pricing (verified on nscale.com/product/inference featured model list), Alibaba Cloud Qwen3 8B / Qwen3 4B Thinking / Qwen3 4B Instruct — open chat models with reasoning variants; per-1M token pricing Nscale 采用充值余额(credit)按量计费,价格按模型分列,在 Nscale 控制台 AI Services → Models 与 /v1/models API 中按每 1M 输入 token 报价(文本模型);图像模型输入按 0 计费。定价结构经 serverless-openapi.yaml ModelPricing schema 验证(input = 每 1M 输入 token 美元价)。 输出按每 1M 输出 token 计费(文本模型);图像模型按每百万像素计费(ModelPricing.output 定义为 per-million output pixels)。具体美元单价在控制台与 /v1/models 端点内按模型展示,注册充值后可查。 ❌ 无中国大陆直连端点。inference.api.nscale.com 是单一全球端点,数据中心集中于欧洲(英国 Loughton、挪威 Glomfjord/Narvik)与美国(德州、西弗吉尼亚);大陆生产流量需代理或中转,跨大西洋/跨太平洋首字节延迟通常 200-350ms。主打 EU 数据主权与在国境内处理,适合欧洲合规场景而非中国区低延迟。 无永久免费档。新用户注册后通过 Google SSO 登录,官方提示早期用户可能获得免费信用额度(promotional offer,非保证)。默认需先充值(credit)才能调用 serverless 推理 API。 查看详情 →
Morph Kimi K3 (2.8T MoE) — flagship Fable-tier open-weight model at ~100 tok/s with 1M context; per-1M input/output token pricing (verified on docs.morphllm.com llms.txt Fast Models table 2026-08-28), GLM-5.2 (744B MoE) — Opus-tier open-weight model with 1M context; OpenAI `service_tier` parameter supported for default/standby processing 按模型分别计费,按每 1M 输入 token 计(文本模型)。从 llms.txt 验证(2026-08-28):Kimi K3 $2.90、Qwen 3.5 397B $0.50、GLM-5.2 744B $1.10、GLM-5.3-Flash $0.15、MiniMax M3 $0.30、MiniMax M2.7 $0.279、DeepSeek V4 Flash (beta) $0.12、Qwen 3.8/3.6 27B $0.289、Gemma 4 31B $0.14。Qwen 3.5 397B 缓存输入 $0.30 per 1M(其他模型无独立缓存价)。Fast Apply 用 `auto` 模型按 prompt+completion 两档计(OpenRouter 验证 morph-v3-fast $0.0008/$0.0012 per 1k、morph-v3-large $0.0009/$0.0019 per 1k),Compact、Reflexes、WarpGrep 与 Router 按事件计费(详见 pricingEN)。 按模型分别计费,按每 1M 输出 token 计(文本模型)。从 llms.txt 验证:Kimi K3 $14.00、Qwen 3.5 397B $3.50、GLM-5.2 744B $4.10、GLM-5.3-Flash $0.50、MiniMax M3 $1.20、MiniMax M2.7 $1.20、DeepSeek V4 Flash (beta) $0.278、Qwen 3.8/3.6 27B $2.40、Gemma 4 31B $0.40。Fast Apply 输出按 morph-v3-fast $1.20 / morph-v3-large $1.90 per 1M token;Fast Apply 用 `auto` 推荐档(5,000-10,500 tok/s, ~98% 精度)按实际路由模型分档计费。Kimi K3 `service_tier: standby` 走 GLM-5.2 同等待遇,标准按 token 计。 ❌ 无中国大陆直连端点。base URL https://api.morphllm.com 为单一全球端点,Y Combinator 背书的美国总部(AutoInfra, Inc.);OpenRouter 上 morph/morph-v3-fast 和 morph/morph-v3-large 走多区域路由,但官方端点不在大陆。中国大陆生产流量需代理或中转,跨太平洋首字节延迟通常 200-400 ms。 无永久免费档(无零元 starter plan)。新用户注册即可获得 API key 并按 pay-as-you-go 充值;官方为初创公司提供 Startup Credits(最高 $5,000),但需单独联系销售申请。Morph 主页明确「Start free. Get API Key」,但实际计费从首次 token 调用开始。 查看详情 →
Firecrawl Scrape (v2 /scrape) — convert any URL into clean markdown / HTML / structured JSON via the JSON-mode schema option; 1 credit per page, Crawl (v2 /crawl) — recursively crawl a website and return content for every linked page; 1 credit per page scraped 纯 credit-based 按 endpoint 与特性消耗,每月按 plan 重置额度;额外 credit 按 $5 一档购买(每档 1k–5k credits 不等)。从 firecrawl.dev/pricing 与 docs.firecrawl.dev/billing.md 验证(2026-08-29):Free 1k/月 $0;Hobby 5k/月 $19(年付 $16);Standard 100k/月 $83(年付 $99);Growth 500k/月 $333(年付 $399);Scale 1M/月 $599;Enterprise Custom。Credit 单价:Scrape / Crawl / Map 1 cr/page、Search 2 cr/10 results、Monitor 1 cr/page/check、Interact 2–7 cr/browser-minute。Batch Scrape / Extract 按子操作同价。Agent 5 daily runs free,超量动态计费。 无独立 output 维度——Firecrawl 输出的是 scraped markdown / HTML / JSON 内容,计费仅按输入端的 credit 消耗与并发浏览器占用。出站数据本身不另收费;唯一例外是 Search 在 ZDR(zero-data-retention)企业档下按企业价另议。 ❌ 无中国大陆直连端点。Firecrawl 总部位于美国旧金山(YC W23),API base URL https://api.firecrawl.dev 单区域部署;docs.firecrawl.dev/billing.md 未列出任何亚太 / 中国路由层。中国大陆生产流量需代理或中转,跨太平洋首字节延迟通常 200-400 ms;Free 档 1000 credits 评估期可代理试用,但稳定流量建议自建代理或自部署 Firecrawl 自托管版(self-hosted Docker / Kubernetes)。 Free 档永久免费:1,000 credits / 月,无需信用卡,支持 Scrape / Crawl / Map / Search / Batch Scrape / JSON mode 全部 endpoints;2 concurrent browsers;50,000 max queued jobs。适合评估与轻量使用,但 credit 重置周期每月初清零、当月用完不再补。Agent 5 daily runs free 适用于所有 plan(含 Free)。 查看详情 →

❓ 常见问题

哪个 AI API 最便宜?

按绝对价格,字节豆包DeepSeek-V3最便宜。但要综合考虑模型能力,DeepSeek在性价比上最强,是目前最推荐的低价高能方案。

国内可以用 OpenAI 和 Claude API 吗?

理论上可以,但需要代理服务。国内直连会被限制或速度极慢。建议使用国内可直连的方案(如 DeepSeek、智谱GLM),或使用API 反代服务

DeepSeek API 真的能用吗?质量如何?

DeepSeek-V3DeepSeek-R1已经达到 GPT-4 级别,在编程和数学推理上甚至更强。DeepSeek-R1 支持思维链输出,推理过程透明。唯一缺点是生态不如 OpenAI 丰富,但价格是 OpenAI 的几十分之一。

如何申请这些 AI API?

每个厂商流程类似:注册账号 → 实名认证(国内平台)→ 充值/购买资源包 → 获取 API Key → 开始调用。具体教程请查看我们的使用教程页面。

📚 最新教程

深度指南与免费额度对比,每周更新

新闻分析

OpenRouter Classifiers:AI API 成本与合规追踪

OpenRouter Classifiers 自动标记每次 API 调用的部门、任务类型和 Agent 复杂度。追踪成本、合规与模型使用。

事件分析

GPT-5.6 Sol 入侵 Hugging Face:API 安全实战复盘

GPT-5.6 Sol 智能体突破 OpenAI 沙箱,10 天内对 Hugging Face 发动自主攻击。API 调用方真正需要的五种防护模式。

AI 安全 智能体沙箱 14 分钟阅读
提供商评测

Weaviate Cloud 2026:开源向量数据库与混合检索评测

Weaviate Cloud 评测:开源向量数据库(Apache 2.0)、混合检索(HNSW+BM25)、8 种内置 Embedding、Query Agent 自然语言查询。免费 10 万对象,Flex $45 起。

向量数据库 混合检索 13 分钟阅读
提供商评测

Portkey 2026:AI Gateway + 可观测性平台评测

Portkey AI Gateway 2026 评测:200+ 模型 OpenAI 兼容接口、可观测性、Guardrails。免费 10K 请求/月,Hobby $49/月,BYOK 零加价。

ai 网关 可观测性 14 分钟阅读
提供商评测

Chroma 2026:Python 优先向量数据库评测

Chroma Cloud 评测:Python 优先向量数据库、21k+ stars、Apache 2.0。免费 50K 向量 + 5K 查询/月,Pro $0.30/M 向量-月。vs Pinecone/Qdrant/Weaviate。

向量数据库 python 13 分钟阅读
提供商评测

Qdrant Cloud 2026:Rust 高性能向量数据库评测

Qdrant Cloud 评测:Rust 编写的开源向量数据库、原生混合检索、命名向量、9 种 FastEmbed、GPU 索引。免费 1GB + 0.5M 向量,Standard $25/月起。

向量数据库 rust 14 分钟阅读
提供商评测

FriendliAI 2026 评测:前沿推理 API 与专属 GPU 端点

按 Token 计费 API 从 $0.14/M 起,专属 GPU $2.9/时起。SOC 2、HIPAA 合规,59 万+ 模型。

推理 前沿模型 12 分钟阅读
热点分析

Claude Opus 5 API:接近 Fable 5 智能,半价即可拥有

Claude Opus 5($5/$25 每百万 Token)以半价达到 CursorBench 上 Fable 5 的 99.5% 性能。ARC-AGI 3 分数是次优模型 3 倍。

价格 基准测试 14 分钟阅读
价格对比

GPT-5 vs Claude 4 vs Gemini 2026:三大 API 价格对比

GPT-5、Claude 4、Gemini 三大 API 价格全面对比,含上下文窗口、功能差异和迁移建议。

价格 对比 8 分钟阅读
提供商评测

Dify 2026 评测:开源 LLM 应用构建平台与 API

Dify(15 万 GitHub stars)是开源 LLM 应用平台。Sandbox 免费、Professional $59、Team $159。RAG 流水线、可视化工作流、100+ 模型。

平台 开源 14 分钟阅读
热点分析

OpenRouter 缓存 + 粘性路由 2026 实测

OpenRouter 缓存读取最低 0.1x(Anthropic/DeepSeek/Qwen),粘性路由锁定 warm 节点。6 轮 Agent 成本对比表 + 4 个 cache miss 排查。

OpenRouter 缓存 粘性路由 11 分钟
平台评测

Pinecone 2026 评测:生产 RAG 托管向量数据库

Pinecone 是 Notion AI、Shopify Sidekick、Cohere 企业 RAG 的底层向量数据库。Serverless $50/月起,$50 免费信用,SOC-2/HIPAA 合规。对比 Weaviate、Qdrant、Milvus、pgvector。

vector 12 分钟
平台评测

Arize Phoenix 2026 评测:开源 LLM 可观测性平台

Arize Phoenix(Elastic-2.0,10.6k 颗星):开源 LLM 可观测性平台,OpenTelemetry 原生 trace,17+ LLM 框架自动接入,10 GiB Cloud 免费层。

OpenTelemetry Elastic-2.0 可自托管
新闻解读

Qwen3.8 vs Kimi K3:两款 2T+ 开源模型 API 对比

Qwen3.8-Max-Preview(阿里 Token Plan)vs Kimi K3(月之暗面 2.8T、1M 上下文):视觉、工具调用、定价对比。核验后的 API 接入与成本测算。

新闻解读

Kimi K3 全面解读:3 万亿级开源旗舰

月之暗面 K3 深度解读:2.8 万亿参数、1M 上下文、原生视觉、Claude Code 集成,权重 7 月 27 日前开源。价格、架构、API、生态集成全梳理。

2.8 万亿参数 1M 上下文 开源 Claude Code
提供商评测

Vercel AI Gateway 2026:零加价 AI 路由层

每月 $5 免费 Credit、Token 零加价(含 BYOK)、AI SDK v5/v6 原生集成。对比 OpenRouter、Cloudflare、Portkey、LiteLLM、Helicone。

Provider Guide

Qwen3.5 API 2026:阿里云百炼 vs 开源权重

Qwen3.5 与阿里云百炼 API 完全指南 2026:核验后的定价、三种访问路径(托管 API、开源权重、聚合器),对比 Kimi K3 与 DeepSeek V4。

Provider Review

RunPod 2026:按 GPU 秒计费的云平台

按 GPU 秒计费从 $0.69 RTX 4090 到 $7.39 B300 288GB;Serverless FlashBoot 亚 200ms 冷启动;31 个全球区域。横向对比 Modal、Baseten、Replicate。

Provider 评测

Helicone 2026 评测:开源 LLM 可观测性平台

Apache-2.0 开源 LLM 可观测性平台(GitHub 5.9k stars)。AI Gateway、HQL 查询语言、SOC-2/HIPAA 合规。Pro $79/月 对比 Portkey/LiteLLM。

🦀
新闻解读

Claude Code 迁移 Bun→Rust:Anthropic 11 天烧 16.5 万美元

Anthropic 用 Claude Code 把 Bun 从 Zig 迁到 Rust:100 万行、100% 测试通过、16.5 万美元 API 费。真实成本、Prompt 缓存、并行 Agent 实操。

查看解读 →
🌙
提供商评测

Kimi K3 API 评测 2026:1M 上下文与官方价格

官方价格、1M 上下文、原生视觉、工具调用、自动缓存、OpenAI 兼容代码,以及更低成本的 K2.7 Code / K2.6 选项。

阅读评测 →
🎙️
提供商评测

AssemblyAI API 评测 2026:Universal-3.5 Pro 转录

Universal-3.5 Pro $0.21/hr、Universal-2 $0.15/hr。免费额度 185 小时/月 pre-recorded + 333 小时/月 streaming。Speaker Diarization、Voice Agent 横评 Deepgram、ElevenLabs、Cartesia。

阅读评测 →
🤝
提供商评测

Together AI API 评测 2026:200+ 开源模型 $0.03/M

200+ 开源模型,$0.03/M 起步价,FlashAttention-4 推理,OpenAI 兼容 API,GPU 集群。

阅读评测 →
🐉
Provider Review

腾讯混元 Hy3:OpenRouter 免费至 7-21

腾讯混元 Hy3(295B MoE / 21B 激活)在 OpenRouter 上免费至 2026-07-21。已验证价格、256K 上下文、与 DeepSeek-V4 Flash 对比、上线策略。

2026-07-14 · 16 min

Provider Review

Modal API 2026:Serverless GPU 云平台

Modal 评测:Python 原生 serverless GPU 云、按 GPU 秒计费、vLLM 一行部署、$30/月免费额度。定价对比 Baseten、Replicate、RunPod。

2026年7月13日 · 14 分钟

🛡️
新闻分析

GPT-Red 2026:OpenAI API 安全指南

GPT-Red 是 OpenAI 内部自动化红队模型,不是公开 API。解读 84% 攻击结果、提示注入、工具权限与 agent 安全。

阅读分析 →
🚀
Provider Review

LiteLLM 2026:开源 AI 网关

LiteLLM 评测:开源 AI 网关,统一 OpenAI 格式接入 100+ LLM。免费自部署、成本跟踪、MCP、Fallback 路由。定价对比 Portkey/OpenRouter。

2026年7月15日 · 15 分钟

🧮
News Analysis

Claude Tokenizer 2026:真实账单计算

Anthropic 新分词器让同一文件多产出 1.36–1.73× tokens。标价仍是 $5/$25,但 Opus 4.8 实际成本 = $7.50/$37.50。

2026年7月15日 · 13 分钟

🚀
新闻分析

GPT-5.6 Sol 对比 Opus 4.8:生产迁移实操

Ploy 把默认 agent 从 Opus 4.8 切到 GPT-5.6 Sol 后成本降 27%、构建时间快 2.2 倍。拆解 4 个工程修复 + CLIProxyAPI。

2026-07-13 · 14 分钟

🆓
入门指南

2026 免费 AI API:14 个平台 + 3 家日常推荐

14 个国内外免费 AI API 平台对比,3 家日常推荐,无需信用卡。

2026-07-06 · 12 分钟

💰
横评对比

2026 AI API 免费额度:26 家厂商对比

三大梯队免费额度横评,覆盖所有主流 AI 提供商。

2026-06-04 · 14 分钟

🏷️
横评对比

2026 最便宜 LLM API 价格

按每 token 经济性排序的最便宜 API 提供商。

2026-06-02 · 11 分钟

🔷
提供商评测

Meta Model API 2026 评测:Muse Spark 1.1 深度解析

Meta 官方 API 全面评测:Muse Spark 1.1 定价、Agentic 工具调用、1M 上下文、搜索增强。

2026-07-12 · 15 分钟

准备好开始使用了吗?

立即对比各平台价格,选择最适合你的 AI 方案