Find the Most Cost-Effective AI Token Plan
Compare real-time pricing from 73 AI providers and 30 frontier/mid/budget models including OpenAI, Claude, Google Gemini, Mistral, xAI Grok, Llama, DeepSeek, and more. Always up to date, completely free.
Global AI APIs
China: VPN/proxy neededThe strongest models for global developers. Mainland China users will need a proxy or VPN to access.
OpenAI
#1OpenAI
New accounts receive $5 in credits valid for 3 months
Anthropic Claude
#2Anthropic Claude
Limited Claude.ai free tier (web only); no API credits; Sonnet 5 limited-time pricing at $2/M until Aug 31, 2026
xAI Grok
#4xAI Grok
Gemini Multimodal
#4Gemini Multimodal
Free tier: Gemini Omni Flash 15 RPM, Veo 3.1 2 generations/day, Imagen 4 10 images/day — no credit card required
Google Gemini
#5Google Gemini
Perplexity AI
#6Perplexity AI
Meta Model API
#6Meta Model API
Free tier: 60 RPM, 2M TPM; free credits on signup for US developers (public preview)
Exa
#7Exa
Permanent $20 signup credit (~2,800 Search calls) + automatic $10 Free Tier monthly credit refill. No credit card required, no minimum spend. Free credits apply to every endpoint.
Mistral AI
#9Mistral AI
Free API credits on signup on La Plateforme (within daily rate limits)
Cohere
#10Cohere
Fireworks AI
#11Fireworks AI
Portkey
#11Portkey
10,000 requests/month; 100K log retention; community Slack support; BYOK free; no credit card required
Black Forest Labs
#11Black Forest Labs
Jina AI
#12Jina AI
Groq
#13Groq
fal.ai
#13fal.ai
Cerebras
#14Cerebras
AI21 Labs (Jamba)
#14AI21 Labs (Jamba)
CoreWeave
#14CoreWeave
No permanent free tier. Serverless Inference bills per token; GPU instances bill hourly (On-Demand) with spot discounts. The 0EM (Zero Egress Migration) program waives egress fees during migration into CoreWeave.
DeepInfra
#15DeepInfra
NVIDIA NIM
#15NVIDIA NIM
Free serverless API endpoints for prototyping (no credit card required, generate API key to start)
Amazon Bedrock
#15Amazon Bedrock
AWS Free Tier offers up to $200 credits for new customers (valid 6 months). No permanent free tier for Bedrock.
Hyperbolic
#15Hyperbolic
No permanent free tier and no long-term lock-in. On-demand GPUs are pay-as-you-go billed per hour, with no quota limits and no minimum commitment to start.
Fireworks AI
#15Fireworks AI
$1 in free serverless credits on signup (pay-per-token, postpaid billing); no permanent free tier. On-demand GPUs and training bill by usage.
Replicate
#16Replicate
SambaNova
#16SambaNova
Nebius
#16Nebius
No setup fee and no monthly subscription. Token Factory is pay-as-you-go per token, with elastic dynamic rate limits that auto-scale up to 20x your base allocation as sustained usage grows - no capacity reservation needed to start.
Anyscale
#17Anyscale
New users get $100 in one-time free credits; small project starter credits ($3-$5) are also available.
DigitalOcean Gradient
#17DigitalOcean Gradient
New accounts get $200 in credits valid for 60 days. There is no permanent free tier; after the credit window, serverless inference bills per token (starting around $0.20 per 1M for the smallest hosted model) and GPU Droplets bill by the hour.
Stability AI
#18Stability AI
Hugging Face
#20Hugging Face
Ideogram
#20Ideogram
Lambda
#20Lambda
No permanent free GPU tier; usage-based billing after signup (GPU instances billed per second). 1-Click Clusters and Superclusters require prepaid / committed terms (2 weeks to 1 year).
Luma AI
#21Luma AI
No permanent free API tier. New users get a small one-time free credit grant on the platform, but sustained API usage requires at least the Plus plan at $30/mo (10,000 credits).
Runway
#22Runway
No free API tier. Minimum $12 for 1,000 credits (~20 Gen-4 Turbo clips). Free web app available but no API access.
TwelveLabs
#22TwelveLabs
Free plan: 600 minutes of free video indexing on signup (cumulative, no credit card required). Full API surface available (Pegasus 1.2 / 1.5, Marengo 3.0, Search, Embed, Analyze & Segment). Index retention is 90 days on Free; upgrading to Developer lifts retention to unlimited.
Mixedbread AI
#22Mixedbread AI
Starter plan grants $5 in one-time credits with no card required; 3 workspace users, 10 stores, 100 requests/minute — enough to evaluate a real corpus. No permanent free tier.
Crusoe
#23Crusoe
No permanent free tier; pay-as-you-go only. Volume discounts via sales contact. Serverless Fine-Tuning starts at $0.40/M tokens for sub-16B models.
Letta
#23Letta
Free $0/mo with up to 3 stateful agents and BYOK; Letta Auto and Skills are throttled — usable for evaluation and small personal-assistant use cases.
Reka
#25Reka
Sign up at platform.reka.ai to obtain an API key; billing is purely pay-as-you-go with no minimum commitment, so teams can prototype with a small top-up. No permanent free tier or monthly free credit.
Aleph Alpha
#28Aleph Alpha
No public free tier — self-serve signup creates an account, but pricing is enterprise contract only
Modal
#28Modal
AssemblyAI
#30AssemblyAI
Together AI
#39Together AI
$1 initial credit on signup (requires payment method)
Nscale
#Nscale
No permanent free tier. New users sign up via Google SSO; the quickstart notes early users may be eligible for free promotional credits (not guaranteed). By default you must top up credit before calling the serverless inference API.
Morph
#41Morph
No permanent free tier (no zero-cost starter plan). New signups get an API key and are billed pay-as-you-go from the first token. Morph offers up to $5,000 in Startup Credits for qualified startups, applied via a separate sales contact form (morphllm.com home page). The homepage CTA 'Start free. Get API Key' refers to signup friction, not to a free quota.
Firecrawl
#42Firecrawl
Permanent Free tier: 1,000 credits per month with no credit card required, full access to Scrape / Crawl / Map / Search / Batch Scrape / JSON mode and every other endpoint; 2 concurrent browsers and a 50,000 max queued jobs ceiling. Designed for evaluation and low-volume use; credits reset at the start of each month and do not roll over. Agent's 5-daily-free runs apply across every plan, including Free.
Global APIs with Partial China Access
China: works but unstableWorks from China for some users; quality and uptime may vary. Check provider status before committing.
Azure OpenAI
#2Azure OpenAI
OpenRouter
#3OpenRouter
Vercel AI Gateway
#8Vercel AI Gateway
$5/month free AI Gateway Credits per team (Free-Tier eligible models only)
Novita AI
#18Novita AI
Baseten
#18Baseten
$30 in inference credits on signup (valid 30 days); custom model hosting free tier: 1 deployed model
Tavily
#18Tavily
Free forever: 1,000 credits/month (rolling 30-day window), no credit card, no feature gating; Search/Extract/Crawl/Map/AI Extraction each have separate quotas (1,000 Search + 500 Crawl); overage returns HTTP 429, no surprise billing
Lepton AI
#19Lepton AI
$5 in credits valid for 30 days (enough to fully evaluate any flagship model)
ElevenLabs
#20ElevenLabs
Deepgram
#21Deepgram
RunPod
#21RunPod
$5 starter credit on signup, no credit card required; per-second billing means spiky workloads never pay for idle
Voyage AI
#22Voyage AI
200M free tokens per account (text embedding + reranker), 50M free tokens for legacy domain models
Weaviate
#24Weaviate
Free forever: 100,000 objects, 1 collection, 7-day backup; 2,000 Embedding requests/day; Query Agent 1,000 req/mo
Writer
#25Writer
14-day free trial, no credit card required (Starter plan)
Qdrant
#26Qdrant
Free forever: 1 GB storage + 0.5M 1015-dim vectors; 2 CPU/0.5 GiB RAM instance; unlimited API requests; community Discord support
Cartesia
#28Cartesia
10,000 characters/month free tier (≈10 minutes), no credit card required
Pinecone
#30Pinecone
FriendliAI
#FriendliAI
No free tier — per-token billing, no minimum
Suno
#40Suno
Free tier: 10 songs per day (non-commercial use only), no credit card required, available immediately on signup
Liquid AI
#40Liquid AI
All 14+ LFM models are free to download, run, and fine-tune under the LFM Open License — including commercially (until your company exceeds $10M annual revenue). Research, education, and non-profit use is always free with no revenue limit.
China-Direct AI APIs
No proxy required in ChinaDirect access from mainland China — best latency for CN users. Quality and ecosystem vary by provider.
DeepSeek
#4DeepSeek
Qwen (Alibaba)
#5Qwen (Alibaba)
1M tokens free credit for new users (90-day validity); Qwen3.5 open weights for self-hosting
Meituan LongCat
#6Meituan LongCat
Stepfun
#6Stepfun
New accounts receive 1,000,000 free tokens (30-day validity from registration)
Alibaba Cloud Bailian
#7Alibaba Cloud Bailian
Baidu ERNIE (文心一言)
#8Baidu ERNIE (文心一言)
Moonshot AI Kimi
#9Moonshot AI Kimi
The ¥15 new-user coupon cannot be used for Kimi K3; recharge required
Dify
#9Dify
Sandbox free tier: 200 GPT-4 calls + 5MB vector storage + 10 docs + 2 workflows + 5,000 API calls/month (30-day logs)
Zhipu AI GLM
#10Zhipu AI GLM
Cloudflare AI Gateway
#10Cloudflare AI Gateway
Tencent Hunyuan
#11Tencent Hunyuan
ByteDance Doubao
#12ByteDance Doubao
LiteLLM
#12LiteLLM
SiliconFlow
#12SiliconFlow
¥1-200 free credits on signup (varies by promo), enough for 1-20M tokens
Arize Phoenix
#12Arize Phoenix
Phoenix Cloud free tier: 10 GiB storage per workspace (no time limit)
Block AI (b.ai)
#13Block AI (b.ai)
Helicone
#13Helicone
FreeModel
#14FreeModel
Credits on registration (amount TBD)
APIKEY.FUN
#15APIKEY.FUN
01.AI Yi (零一万物)
#1601.AI Yi (零一万物)
Free rate limit tier on registration (limited RPM/TPM), no free initial credits
MiniMax (Hailuo AI)
#17MiniMax (Hailuo AI)
Mem0
#18Mem0
Hobby free tier: 10,000 memories + 1,000 retrieval API calls/month (unlimited end users), community support
Kling AI (Kuaishou)
#19Kling AI (Kuaishou)
New users get a small one-time free credit grant on the platform to try video/image generation; there is no permanent free API tier — API usage is billed per generated second (video) or per image.
Chroma
#22Chroma
Chroma OSS permanently free, Apache 2.0, no feature limits, no vector count cap (limited by local hardware); Chroma Cloud Free forever: 50K vectors + 5K queries/mo + 0.5 GB storage + 1 Collection + community support
fal.ai
#27fal.ai
No free API tier (prepay $10+ for credit balance); no monthly fees, no subscriptions, no hidden fees; serverless auto-scales to zero when idle
Hume AI
#28Hume AI
Free credits on signup (amount per account), enough to trial EVI 3
AI API Gateways & Aggregators
One API key → 400+ models. Unified access with auto-routing and model flexibility.
OpenRouter
#3300+ models
Gateway to 400+ models from 70+ providers through a single OpenAI-compatible API endpoint
Vercel AI Gateway
#8200+ models
$5/mo free credit, zero markup on tokens (incl. BYOK)
$5/month free AI Gateway Credits per team (Free-Tier eligible models only)
Dify
#9100+ models
GitHub 150,010 stars — the de-facto open-source LLM app platform (langgenius/dify, Apache-style license)
Sandbox free tier: 200 GPT-4 calls + 5MB vector storage + 10 docs + 2 workflows + 5,000 API calls/month (30-day logs)
Cloudflare AI Gateway
#10500+ models
Unified API gateway for 20+ AI providers
Portkey
#11200+ models
🎯 AI Gateway + full-stack observability in one console: Gateway, Logs, Feedback, Guardrails
10,000 requests/month; 100K log retention; community Slack support; BYOK free; no credit card required
LiteLLM
#12100+ models
100+ provider unified interface, popular open-source
SiliconFlow
#12100+ models
All-in-one model hosting in China, 100+ open-source models online
¥1-200 free credits on signup (varies by promo), enough for 1-20M tokens
Arize Phoenix
#12100+ models
Only Elastic-2.0 open-source (10,600+ stars) OpenTelemetry-native LLM observability platform
Phoenix Cloud free tier: 10 GiB storage per workspace (no time limit)
Block AI (b.ai)
#13150+ models
One key for GPT-4o / Claude / Gemini overseas models
fal.ai
#13200+ models
200+ models under one API: FLUX, Kling, Hunyuan, MiniMax, Wan all routed through fal
Helicone
#13100+ models
YC W23, Apache-2.0 open-source LLM observability platform
FreeModel
#1450+ models
Official DeepSeek partner — direct access to DeepSeek V3, R1 models
Credits on registration (amount TBD)
APIKEY.FUN
#1550+ models
Direct China access to overseas models
Replicate
#1650000+ models
50,000+ open-source models across all AI domains
Novita AI
#1850+ models
Competitive Llama 3.3 70B pricing
Baseten
#1835+ models
Per-second GPU billing, the most transparent pay-as-you-go among Replicate-style competitors
$30 in inference credits on signup (valid 30 days); custom model hosting free tier: 1 deployed model
Mem0
#1850+ models
De-facto AI Agent memory layer (61,493 GitHub stars, Apache-2.0, mem0ai/mem0)
Hobby free tier: 10,000 memories + 1,000 retrieval API calls/month (unlimited end users), community support
Tavily
#187+ models
🎯 Default search API for AI agents: LangChain TavilySearchResults, LlamaIndex TavilyToolSpec, AutoGen/CrewAI/smolagents all default to Tavily
Free forever: 1,000 credits/month (rolling 30-day window), no credit card, no feature gating; Search/Extract/Crawl/Map/AI Extraction each have separate quotas (1,000 Search + 500 Crawl); overage returns HTTP 429, no surprise billing
Hugging Face
#20200000+ models
World's largest AI model community, 200,000+ models
fal.ai
#271396+ models
🧠 1,396+ open-source/commercial models: FLUX.2 Pro, GPT Image 2, Kling Video, Nano Banana 2, ElevenLabs TTS, MiniMax and more
No free API tier (prepay $10+ for credit balance); no monthly fees, no subscriptions, no hidden fees; serverless auto-scales to zero when idle
💰 Token Price Comparison
Sorted: global APIs first, then global-with-partial-China, then China-direct.
| Provider | Popular Models | Input | Output | China Access | Free Tier | |
|---|---|---|---|---|---|---|
| OpenAI | GPT-5, GPT-5-mini | GPT-5: $1.25/M, GPT-5-mini: $0.25/M, GPT-4o: $2.50/M, GPT-4o-mini: $0.15/M | GPT-5: $10/M, GPT-5-mini: $2/M, GPT-4o: $10/M, GPT-4o-mini: $0.60/M | ❌ Proxy required (export control restrictions for China users) | New accounts receive $5 in credits valid for 3 months | View Details → |
| Anthropic Claude | Claude Sonnet 5, Claude Opus 5 | Sonnet 5 (intro): $2/M, Opus 5: $5/M, Sonnet 5/4.5: $3/M, Opus 4.8: $5/M, Haiku 4.5: $0.80/M | Sonnet 5 (intro): $10/M, Opus 5: $25/M, Sonnet 5/4.5: $15/M, Opus 4.8: $25/M, Haiku 4.5: $4/M | ❌ Proxy required (US export controls) | Limited Claude.ai free tier (web only); no API credits; Sonnet 5 limited-time pricing at $2/M until Aug 31, 2026 | View Details → |
| xAI Grok | Grok 4.6, Grok 4.6 Fast | Grok 4.6: $2/M tokens (fast variant $4/M) | Grok 4.6: $6/M tokens (fast variant $12/M) | ❌ Proxy required | View Details → | |
| Gemini Multimodal | Gemini Omni Flash (multimodal, 1M context, audio+video+text), Nano Banana 2 Lite (image generation, $0.034/1K tokens) | Omni Flash: $0.15/M input tokens (multimodal); Nano Banana 2 Lite: $0.034/1K image tokens; Veo 3.1: $0.10/sec video | Omni Flash: $0.60/M output tokens (multimodal); Nano Banana Pro: $0.12/1K image tokens | ❌ Proxy required (Google AI Studio not directly accessible from mainland China; requires stable proxy) | Free tier: Gemini Omni Flash 15 RPM, Veo 3.1 2 generations/day, Imagen 4 10 images/day — no credit card required | View Details → |
| Google Gemini | Gemini 2.5 Pro, Gemini 2.5 Flash | 2.5 Pro: $1.25-10/M, 2.5 Flash: $0.15-1.25/M, 2.0 Flash: $0.10/M | 2.5 Pro: $5-40/M, 2.5 Flash: $0.30-5/M, 2.0 Flash: $0.40/M | ❌ Proxy required (Google Cloud not accessible in CN) | View Details → | |
| Perplexity AI | Sonar Pro, Sonar | Sonar Pro: $3/M, Sonar: $1/M | Sonar Pro: $15/M, Sonar: $1/M | ❌ Proxy required | View Details → | |
| Meta Model API | Muse Spark 1.1 (muse-spark-1.1) | $1.25/1M tokens | $4.25/1M tokens | ❌ Proxy required (Meta Model API public preview for US developers only; China access requires proxy) | Free tier: 60 RPM, 2M TPM; free credits on signup for US developers (public preview) | View Details → |
| Exa | Search — neural semantic search across the web; 6 search types (auto / instant / fast / deep-lite / deep / deep-reasoning); returns 10 results by default, $1/1k per extra result above 10; AI page summaries $1/1k pages, Contents — full page text / highlights / summaries for known URLs; $1/1k pages per content type (text, highlights, summary billed separately) | Pure pay-as-you-go, billed per endpoint and search-type. Search $7/1k requests (base 10 results) + $1/1k per extra result above 10 + $1/1k AI summaries; Contents $1/1k pages; Answer $5/1k; Monitors $15/1k; Deep Search $12-15/1k; Agent fixed effort $0.012-$1.00 per request or usage-based $0.10/ACU + tool calls (default $5 auto cap, $20 max cap). New accounts get $20 in free credits (~2,800 searches); Free Tier adds $10/month. | Same as input. Enterprise custom volume + Zero Data Retention + SLA + postpaid invoice billing. | ❌ Proxy required for mainland China. Exa runs the primary API at https://api.exa.ai from a single-region US deployment. There is no documented ICP-registered mainland China endpoint as of 2026-08-30. Mainland China production traffic typically needs a proxy or relay; transpacific first-byte latency is usually 100-300 ms. Enterprise plans can negotiate custom routing and dedicated infrastructure for regulated workloads (HIPAA / GDPR data residency). | Permanent $20 signup credit (~2,800 Search calls) + automatic $10 Free Tier monthly credit refill. No credit card required, no minimum spend. Free credits apply to every endpoint. | View Details → |
| Mistral AI | Mistral Large 2, Mistral Small | Large 2: $2/M, Codestral: $1/M, Ministral 8B: $0.10/M | Large 2: $6/M, Codestral: $3/M, Ministral 8B: $0.10/M | ❌ Proxy required | Free API credits on signup on La Plateforme (within daily rate limits) | View Details → |
| Cohere | Command R+, Command R7 | Command R7: $2.50/M, Command R: $0.50/M, Embed: $0.10/M | Command R7: $10/M, Command R: $1.50/M | ❌ Proxy required | View Details → | |
| Fireworks AI | Llama 3.3 70B, Firefunction-v2 | Firefunction-v2: $0.90/M, Llama 3.3 70B: $0.90/M, DeepSeek-R1: $2.00/M | Firefunction-v2: $0.90/M, Llama 3.3 70B: $0.90/M, DeepSeek-R1: $2.00/M | ❌ Proxy required | View Details → | |
| Portkey | OpenAI GPT-4o / GPT-5 / GPT-5.6 family, Anthropic Claude Opus 5 / Sonnet 5 / Haiku | Free tier: 10,000 requests/month; Hobby $49/month (100K requests); Growth $249/month (1M requests); Enterprise contract | Per-request billing, no token routing markup; BYOK free; logs/observability/caching gated by subscription tier | ❌ Proxy required (portkey.ai is unstable from mainland China; self-hosted open-source edition recommended for CN teams) | 10,000 requests/month; 100K log retention; community Slack support; BYOK free; no credit card required | View Details → |
| Black Forest Labs | FLUX.2 [pro], FLUX.1.1 [pro] | FLUX.2 [pro]: $0.05/MP, FLUX.1.1 [pro]: $0.04/MP, FLUX.1 [schnell]: $0.003/MP, FLUX.1 Fill: $0.05/MP | Flat per-megapixel billing, no tier-based markup | ❌ Proxy required (AWS US-East + EU-Frankfurt deployment) | View Details → | |
| Jina AI | jina-embeddings-v3, jina-embeddings-v2-base-en | Embeddings v3: $0.02/M tokens (10K free); Reranker v3: $0.018/M tokens (10K free); Reader: $0.02/M tokens (free up to 1M tokens); CLIP-v1: $0.02/M tokens | Flat per-million-token billing, no tier-based markup; batch discount of 50% available on all endpoints | ❌ Proxy required (AWS US-East + EU-Frankfurt deployment) | View Details → | |
| Groq | Llama 3.3 70B, Llama 3.1 405B | Llama 3.3 70B: $0.59/M, DeepSeek-R1: $0.89/M, Mixtral 8x7B: $0.15/M | Llama 3.3 70B: $0.79/M, DeepSeek-R1: $0.89/M, Mixtral 8x7B: $0.15/M | ❌ Proxy required | View Details → | |
| fal.ai | FLUX.2 [pro] / [dev] / [schnell], Kling Video 2.1 / 2.0 | FLUX.2 [pro]: $0.05/MP (image), $0.08/sec (video); Kling 2.1: $0.10/sec; HunyuanVideo 1.5: $0.08/sec | Flat per-second / per-megapixel billing, no tier-based markup | ❌ Proxy required (AWS US-East deployment) | View Details → | |
| Cerebras | Cerebras Llama 3.3 70B, Cerebras Llama 3.1 405B | $0.10/M to $0.60/M tokens(所有模型统一 $0.60/M input) | $0.10/M to $0.60/M tokens(统一 $0.60/M output) | ❌ Proxy required | View Details → | |
| AI21 Labs (Jamba) | Jamba 1.5 Mini, Jamba 1.5 Large | Jamba 1.5 Mini: $0.20/M tokens, Jamba 1.5 Large: $2/M tokens | Jamba 1.5 Mini: $0.20/M tokens, Jamba 1.5 Large: $2/M tokens | ❌ Proxy required | View Details → | |
| CoreWeave | NVIDIA HGX H100 ($49.24/hr), HGX H200 ($50.44/hr), HGX B200 ($68.80/hr), A100 80GB ($21.60/hr), L40S ($18.00/hr), L40 ($10.00/hr), GB200 NVL72 ($42.00/hr) — On-Demand/Spot GPU instances, Serverless Inference (W&B Inference, pay-per-token): GLM 5.2, Kimi K2.6/K2.7, DeepSeek R1/V3, Llama 3.x, Qwen — OpenAI-compatible API | GPU On-Demand hourly: NVIDIA HGX H100 $49.24, HGX H200 $50.44, HGX B200 $68.80, A100 80GB $21.60, L40S $18.00, L40 $10.00, GB200 NVL72 $42.00. Spot discounts: H100 $19.71, H200 $20.93, B200 $34.11, A100 $9.65. Serverless Inference bills per token via W&B Inference; Dedicated Inference bills on GPU-hour or node basis. | ❌ No mainland China direct endpoint. Data centers concentrated in the US; expanded to Indonesia (first APAC region), UK (two operational data centers), and Sweden in 2026. China access requires proxy; cross-Pacific latency is high. | No permanent free tier. Serverless Inference bills per token; GPU instances bill hourly (On-Demand) with spot discounts. The 0EM (Zero Egress Migration) program waives egress fees during migration into CoreWeave. | View Details → | |
| DeepInfra | Meta-Llama-3.3-70B-Instruct, Meta-Llama-3.1-405B-Instruct | Llama 3.3 70B: $0.49/M, DeepSeek-V3: $0.99/M, DeepSeek-R1: $1.09/M | Llama 3.3 70B: $0.73/M, DeepSeek-V3: $0.99/M, DeepSeek-R1: $1.09/M | ❌ Proxy required | View Details → | |
| NVIDIA NIM | Nemotron-3 Ultra 550B (a55B), Nemotron-3 Super 120B (a12B) | Partner-dependent (Nemotron-3 Ultra 550B: $0.50-0.90/1M tokens) | Partner-dependent (Nemotron-3 Ultra 550B: $1.70-3.60/1M tokens) | ❌ Proxy required (US-based platform, export control restrictions apply) | Free serverless API endpoints for prototyping (no credit card required, generate API key to start) | View Details → |
| Amazon Bedrock | Claude Opus 4.8 / 4.7 / 4.6 / 4.5, Claude Sonnet 4.6 / 4.5 / 4 | Claude Sonnet 4.5: $6/M, DeepSeek V3.2: $0.62/M, Mistral Large 3: $0.50/M, GPT-5.5: $5.50/M, GPT-5.4: $2.75/M, Grok 4.3: $1.25/M, Qwen3 235B: $0.23/M, Kimi K2.5: $0.60/M | Claude Sonnet 4.5: $30/M, DeepSeek V3.2: $1.85/M, Mistral Large 3: $1.50/M, GPT-5.5: $33/M, GPT-5.4: $16.50/M, Grok 4.3: $2.50/M, Qwen3 235B: $0.91/M, Kimi K2.5: $3/M | ❌ Proxy/VPN required for mainland China users (Bedrock not available in AWS China regions) | AWS Free Tier offers up to $200 credits for new customers (valid 6 months). No permanent free tier for Bedrock. | View Details → |
| Hyperbolic | DeepSeek R1, DeepSeek R1-0528 | Per-token inference scaled from market GPU rates; H100 from $2.89/GPU-hour, H200 $3.49/GPU-hour | Pay-as-you-go GPU compute billed per hour; reserved discounts for committed capacity | ❌ Proxy required | No permanent free tier and no long-term lock-in. On-demand GPUs are pay-as-you-go billed per hour, with no quota limits and no minimum commitment to start. | View Details → |
| Fireworks AI | DeepSeek V4 Pro — $1.74 input / $0.145 cached / $3.48 output per 1M tokens (Serverless Standard; Priority $2.61 / $0.218 / $5.22), DeepSeek V4 Flash (0731) — $0.22 input / $0.007 cached / $0.66 output per 1M (Standard) | Serverless (Standard, $/1M tokens): DeepSeek V4 Pro $1.74, V4 Flash $0.22, Kimi K3 $3.00, K2.7 Code $0.95, MiniMax M3 $0.30, Qwen 3.7 Plus $0.40, GPT OSS 120B $0.15, GLM 5.1 $1.40. Embeddings $0.008-$0.016/1M by tier. On-demand GPUs per hour: H100 $7.00, H200 $7.00, B200 $10.00, B300 $12.00, GB300 $18.00 (rising Sep 1). | Serverless output ($/1M): DeepSeek V4 Pro $3.48, V4 Flash $0.66, Kimi K3 $15.00, MiniMax M3 $1.20, Qwen 3.7 Plus $1.60, GPT OSS 120B $0.60, GLM 5.1 $4.40. Cached input far lower (DeepSeek V4 Pro $0.145). Training: SFT $0.50/1M train tokens (≤16B) to $10.00 (>300B); Serverless Training API bills tokens only, no idle GPU cost. | ❌ No mainland China direct endpoint. As a US-based platform, direct access from the mainland requires proxy or relay; cross-Pacific first-byte latency typically 150-250ms. Closest APAC endpoint is the Tokyo multi-region. | $1 in free serverless credits on signup (pay-per-token, postpaid billing); no permanent free tier. On-demand GPUs and training bill by usage. | View Details → |
| Replicate | Llama 3.3 70B, DeepSeek-R1 | 模型调用按运行时长 + GPU 型号计费,$0.00025/s/A100起步 | 按运行时长计费(非 token 计费模式) | ❌ Proxy required | View Details → | |
| SambaNova | Llama 3.3 70B, Llama 3.1 405B | Llama 3.3 70B: $0.50/M, DeepSeek-R1: $0.80/M, Qwen 72B: $0.40/M | Llama 3.3 70B: $0.50/M, DeepSeek-R1: $0.80/M, Qwen 72B: $0.40/M | ❌ Proxy required | View Details → | |
| Nebius | DeepSeek V4 Flash, DeepSeek V4 Pro | Per-token OpenAI-compatible inference: DeepSeek-V4-Flash $0.14/M input, $0.28/M output; Kimi K3 $3/$15; MiniMax M3 $0.30/$1.20; Llama 3.3 70B $0.13/$0.40 | Two flavors per model - Base and Fast (-fast suffix) - same outputs, Fast trades higher token price for lower latency via speculative decoding | ❌ Proxy required | No setup fee and no monthly subscription. Token Factory is pay-as-you-go per token, with elastic dynamic rate limits that auto-scale up to 20x your base allocation as sustained usage grows - no capacity reservation needed to start. | View Details → |
| Anyscale | Llama 3.3 70B, Llama 3.1 405B | Llama 3.3 70B: $0.39/M, DeepSeek-R1: $1/M, Mixtral 8x22B: $0.90/M | Llama 3.3 70B: $0.39/M, DeepSeek-R1: $1/M, Mixtral 8x22B: $0.90/M | ❌ Proxy required | New users get $100 in one-time free credits; small project starter credits ($3-$5) are also available. | View Details → |
| DigitalOcean Gradient | Llama 3.3 70B, Llama 3.1 405B | Llama 3.3 70B: $0.50/M, DeepSeek-R1: $0.80/M | Llama 3.3 70B: $0.50/M, DeepSeek-R1: $0.80/M | ❌ Proxy required | New accounts get $200 in credits valid for 60 days. There is no permanent free tier; after the credit window, serverless inference bills per token (starting around $0.20 per 1M for the smallest hosted model) and GPU Droplets bill by the hour. | View Details → |
| Stability AI | Stable Diffusion 3.5, Stable Diffusion XL | SD3.5: $0.0065/image, SDXL: $0.01/image, Stable Code: $0.50/M tokens | 按图像/视频/音频输出计费,非 token 计费 | ❌ Proxy required | View Details → | |
| Hugging Face | Llama 3.3 70B, Llama 3.1 405B | 推理 API: $0.04-0.60/M tokens 取决于模型 | 推理 API: $0.04-0.60/M tokens | ❌ Proxy required (huggingface.co blocked in China) | View Details → | |
| Ideogram | Ideogram 3.0, Ideogram Turbo | Ideogram 3.0: $0.04/image (standard), $0.08/image (Turbo) | 4 images per generation (default), additional images via paid credits | ❌ Proxy required | View Details → | |
| Lambda | NVIDIA HGX B200 180GB (on-demand GPU instances; $6.69/GPU/hr), NVIDIA H100 SXM 80GB ($3.99/GPU/hr on-demand) | GPU instance on-demand per GPU/hr: NVIDIA B200 SXM6 $6.69, H100 SXM $3.99, A100 SXM 80GB $2.79, A100 SXM 40GB $1.99, GH200 $2.29, A100 PCIe $1.99, A6000 $1.09, A10 $1.29, V100 $0.79. 1-Click Cluster reserved per GPU/hr: B200 $9.86 (16 GPU) down to $8.87 (256+), H100 $6.16 down to $5.54. Superclusters / Private Cloud (4000+ GPUs) via sales. No egress fees. | Billed by the second per GPU instance (no idle GPU cost); clusters and reserved capacity carry a minimum commitment (2 weeks to 1 year). Managed orchestration: Managed Kubernetes and Slurm included at no extra per-node fee; Lambda Stack one-line install. No per-token inference pricing — Lambda's Inference API is winding down; self-hosted inference billed as GPU runtime. | ❌ No mainland China direct endpoint; data centers in the US (California and other regions), China access requires proxy or overseas relay. | No permanent free GPU tier; usage-based billing after signup (GPU instances billed per second). 1-Click Clusters and Superclusters require prepaid / committed terms (2 weeks to 1 year). | View Details → |
| Luma AI | Ray 3.2 (flagship video, 1080p, multi-keyframe up to 16 frames, V2V up to 20s), Ray 3.14 (previous-gen video model) | Credit-based subscription. Plans: Plus $30/mo (10,000 credits), Pro $90/mo (40,000 credits), Ultra $300/mo (150,000 credits). Ray 3.2 video: 1080p text-to-video 400 credits/5s, 720p 100 credits/5s, Draft 20 credits/5s; Seedance 2.0 1080p 240 credits/sec, 4K 959 credits/sec. | Image credits: Uni-1 30 credits/image, Seedream 1-3 credits/image, GPT Image 2 from 3 (Low-1K) to 255 (High-4K) credits. Audio: ElevenLabs v3 TTS 21 credits/1,000 chars. Utilities: background removal 1 credit/image, reframe video 32 credits/sec, upscale to 4K 17 credits/sec. | ❌ Proxy required (no mainland China direct access; API served from global AWS/GCP edge regions) | No permanent free API tier. New users get a small one-time free credit grant on the platform, but sustained API usage requires at least the Plus plan at $30/mo (10,000 credits). | View Details → |
| Runway | Gen-4 Turbo, Gen-4 Alpha | Gen-4 Turbo: 50 credits/5s clip (~$0.50), Gen-4 Alpha: 100 credits/5s (~$1.00), Act-Two: 150 credits/5s (~$1.50), Frames: 200 credits/sequence (~$2.00) | Credit-based system (~$0.01/credit). Credit packs: $12 (1,000 credits) to $200 (25,000 credits). Enterprise: custom pricing. | ❌ Proxy required (AWS US-East + GCP US-Central hosting) | No free API tier. Minimum $12 for 1,000 credits (~20 Gen-4 Turbo clips). Free web app available but no API access. | View Details → |
| TwelveLabs | Pegasus 1.5 (flagship multimodal video understanding: frames + audio + speech + on-screen text), Pegasus 1.2 (previous-gen, lower-cost option: frames + audio + speech) | Free plan: 600 minutes of video indexing at no cost. Pegasus 1.2 video indexing $0.042/min (one-time), input video $0.021/min, output text $0.0075/1k tokens, embedding infrastructure $0.0015/indexed minute/month. Developer plan upgrades indexes past 90 days and unlocks higher rate limits. | Identical per-minute and per-1k-tokens billing. Marengo embeddings billed by indexed minute. No subscription fee on Developer plan beyond per-usage costs. | ❌ Proxy required (no mainland China direct endpoint; API hosted on AWS us-west-2 / us-east-1; SDKs available in Python / Node / Go / Java) | Free plan: 600 minutes of free video indexing on signup (cumulative, no credit card required). Full API surface available (Pegasus 1.2 / 1.5, Marengo 3.0, Search, Embed, Analyze & Segment). Index retention is 90 days on Free; upgrading to Developer lifts retention to unlimited. | View Details → |
| Mixedbread AI | Toast 1 — specialized search model (2026-08-13): matches/outperforms Claude Opus 5 and GPT-5.6 Sol on knowledge work, up to 10x cheaper and 12x faster; input $0.50 / cached $0.06 / output $1.20 per 1M LLM tokens (launch $0.30 / $0.036 / $0.72), mxbai-embed-large-v1 — 1,024-dimension open embedding model (Apache-2.0), 50M+ downloads across the mxbai family | Usage-based across three buckets: Indexing (Fast $1.50/1M content tokens; High Quality with OCR/transcription/summaries/multimodal enrichment $3/1M), Search (Grep $0.10/1K queries, Semantic $4/1K, Toast 1 $1/1K; rerank adds $3.50 or $1.50/1K), and Storage $0.50/1M content tokens per month. Toast 1 also bills input $0.50/1M LLM tokens (launch $0.30). | Toast 1 output $1.20/1M LLM tokens (launch $0.72, 40% off); cached input $0.06/1M (launch $0.036), cache writes free. Semantic search rerank adds $3.50/1K queries. | ❌ No mainland China direct endpoint. Mixedbread is a US/Europe-based platform (legal entity mixedbread ai inc., Berlin/San Francisco), so direct access from the mainland requires a proxy or relay, with cross-Pacific first-byte latency typically 150-250ms. | Starter plan grants $5 in one-time credits with no card required; 3 workspace users, 10 stores, 100 requests/minute — enough to evaluate a real corpus. No permanent free tier. | View Details → |
| Crusoe | DeepSeek V3 0324 (legacy frontier: $0.50 input / $1.50 output per 1M tokens), DeepSeek V4 Flash (efficient frontier: $0.14 input / $0.28 output per 1M tokens, 70B tier) | Pay-as-you-go per 1 million tokens, four parameter-count tiers: <16B ($0.40) / 16B-70B ($2.50) / 70B-300B ($6.00) / >300B ($10.00). Serverless Fine-Tuning uses the identical 4-tier structure. Managed Inference covers DeepSeek V3 0324, V4 Pro, V4 Flash, GLM 5.1/5.2, GPT-OSS 20B/120B, Gemma 4 31B-it, Kimi K2.6, Llama 3.1 8B / 3.3 70B, Nemotron 3 family (VoiceChat/Ultra 550B/3.5 Lightning/3 Nano/3 Super/Nano Omni 30B), Qwen3 8B/235B A22B Instruct 2507/Qwen3.5 2B+9B/Qwen3.6 35B A3B, Yutori n1.5. Cached tokens billed at $0.03-$1.50 per 1M. | Self-Serve Deployments: NVIDIA H100 80GB HGX $5.50/hr, NVIDIA H200 141GB HGX $6.00/hr (dedicated endpoints for open and fine-tuned models). Tailored Deployments and Provisioned Throughput negotiated via sales. Managed Kubernetes $0.10/cluster-hour; Container Registry $0.10/GiB-month; Object Storage $0.06/GiB-month. | ❌ No mainland China direct endpoint. Managed Inference API hosted at Crusoe's US data centers (Colorado, Texas, and other US locations); China access requires proxy. | No permanent free tier; pay-as-you-go only. Volume discounts via sales contact. Serverless Fine-Tuning starts at $0.40/M tokens for sub-16B models. | View Details → |
| Letta | Letta Agents SDK — open-source stateful agent runtime (Apache-2.0), with first-class persistent memory (core/archival/recall blocks) and runtime memory editing, Letta Code (`@letta-ai/letta-code`, Node.js 22.19+) — memory-first CLI coding agent, #1 on Terminal-Bench (Dec 2025 launch) | Usage-based: Free $0 (up to 3 stateful agents, BYOK), Pro $20/mo (up to 20 agents, Letta Auto weekly + monthly quota + pay-as-you-go overage), Teams Pro per seat (shared agents + permissions), Developer Plan API-key billing with pure credit usage (no agent cap), Enterprise sales-quoted (self-hosting / custom models). Server-side tools bill CPU time at $0.00015/sec. | Letta Auto usage above the Pro plan's included quota falls back to pay-as-you-go at the model's listed rate (see platform.letta.com/models). Mods, Skills, and Conversations are included in plans, not billed separately. Remote MCP tools run on the MCP provider — no Letta credit cost. | ❌ No mainland China direct endpoint. Letta, Inc. is registered in San Francisco and platform.letta.com is a single global endpoint, so production traffic from mainland China requires a proxy or relay (typical cross-Pacific first-byte latency ~150-250ms). The agent harness is fully open-source (Apache-2.0), so self-hosting on your own infrastructure is a viable workaround. | Free $0/mo with up to 3 stateful agents and BYOK; Letta Auto and Skills are throttled — usable for evaluation and small personal-assistant use cases. | View Details → |
| Reka | Reka Core — top-tier multimodal model (text + image + audio + video); input $2.00 / output $6.00 per 1M tokens, image $0.02 / video $0.08 / audio $0.02 per minute, Reka Flash — cost-efficient model for everyday tasks (text + image + audio + video); input $0.80 / output $2.00 per 1M tokens, image $0.01 / video $0.06 / audio $0.015 per minute | Reka Edge $0.10/M input; Reka Flash $0.80/M; Reka Core $2.00/M; Reka Flash Research bills per 1K requests ($25 standard, $35 parallel-low, $60 parallel-high). Pure usage-based with no minimum commitment. | Reka Edge $0.10/M; Reka Flash $2.00/M; Reka Core $6.00/M. Multimodal usage (image, video, audio) bills separately per minute: Flash video $0.06/min, Core $0.08/min; images billed per call (Flash $0.01, Core $0.02). | ❌ No mainland China direct endpoint. Reka is registered at 530 Lawrence Expressway, Sunnyvale, CA; api.reka.ai is a single global endpoint, so production traffic from mainland China requires a proxy or relay (typical cross-Pacific first-byte latency ~150-250ms). The open weights can be self-hosted on NVIDIA GPUs with ≥24GB VRAM via vLLM to bypass direct-access limits. | Sign up at platform.reka.ai to obtain an API key; billing is purely pay-as-you-go with no minimum commitment, so teams can prototype with a small top-up. No permanent free tier or monthly free credit. | View Details → |
| Aleph Alpha | Pharia-1-LLM-7B-control, Pharia-1-LLM-7B-control-aligned | Enterprise contract — no public per-token price (quote-based; PhariaAI on-prem deployment custom-priced) | Same — enterprise quote, no public per-token list | ❌ Not applicable (German sovereign-AI vendor, subject to EU export controls) | No public free tier — self-serve signup creates an account, but pricing is enterprise contract only | View Details → |
| Modal | B300 / B200 / H200 / H100 / A100 / L40S / A10 / L4 / T4 (GPU), Self-hosted LLM inference (vLLM, SGLang, TensorRT-LLM) | Per GPU-second: H100 $0.001097/s, A100 80GB $0.000694/s, T4 $0.000164/s | Same as input (per GPU-second, no idle fees) | ❌ Proxy required (infrastructure on AWS/GCP overseas regions; direct CN access unreliable) | View Details → | |
| AssemblyAI | Universal-3.5 Pro (Async, 18 languages, native code switching, best diarization), Universal-2 (Async, 99 languages, balanced accuracy) | Per audio hour: Universal-3.5 Pro $0.21/hr, Universal-2 $0.15/hr (pre-recorded) | N/A (speech-to-text billed on audio duration, no token concept) | ❌ Proxy required (US-based infrastructure, CN direct access unreliable) | View Details → | |
| Together AI | DeepSeek V4 Pro, Qwen3.6-Plus / Qwen3.7-Max | $0.03-$1.74/MTok (model-dependent) | $0.12-$4.50/MTok (model-dependent) | ❌ Proxy required (US-based service) | $1 initial credit on signup (requires payment method) | View Details → |
| Nscale | Moonshot AI Kimi K2.5 — open chat model via serverless inference; per-1M input/output token pricing (verified on nscale.com/product/inference featured model list), Alibaba Cloud Qwen3 8B / Qwen3 4B Thinking / Qwen3 4B Instruct — open chat models with reasoning variants; per-1M token pricing | Nscale uses credit-based pay-as-you-go billing with per-model rates quoted per 1M input tokens (text models) in the Nscale Console (AI Services → Models) and via the /v1/models API endpoint; image models bill input as 0. The pricing structure is verified against the ModelPricing schema in serverless-openapi.yaml (input = USD per 1M input tokens). | Output bills per 1M output tokens (text models); image models bill per million output pixels (ModelPricing.output defined as per-million output pixels). Exact per-model USD rates are shown in the Console and /v1/models once you top up credit. | ❌ No mainland China direct endpoint. inference.api.nscale.com is a single global endpoint; data centers concentrate in Europe (Loughton UK, Glomfjord/Narvik Norway) and the US (Texas, West Virginia), so China production traffic needs a proxy and transatlantic/transpacific first-byte latency is typically 200-350ms. The platform emphasizes EU data sovereignty and in-country processing, which suits European compliance workloads more than China-region low latency. | No permanent free tier. New users sign up via Google SSO; the quickstart notes early users may be eligible for free promotional credits (not guaranteed). By default you must top up credit before calling the serverless inference API. | View Details → |
| Morph | Kimi K3 (2.8T MoE) — flagship Fable-tier open-weight model at ~100 tok/s with 1M context; per-1M input/output token pricing (verified on docs.morphllm.com llms.txt Fast Models table 2026-08-28), GLM-5.2 (744B MoE) — Opus-tier open-weight model with 1M context; OpenAI `service_tier` parameter supported for default/standby processing | Per-model billing quoted per 1M input tokens (text models). Verified against docs.morphllm.com llms.txt Fast Models table (2026-08-28): Kimi K3 $2.90, Qwen 3.5 397B $0.50, GLM-5.2 744B $1.10, GLM-5.3-Flash $0.15, MiniMax M3 $0.30, MiniMax M2.7 $0.279, DeepSeek V4 Flash (beta) $0.12, Qwen 3.8/3.6 27B $0.289, Gemma 4 31B $0.14. Qwen 3.5 397B cached input bills at $0.30 per 1M; other models have no separate cache rate. Fast Apply via the `auto` model bills at OpenRouter-verified rates (morph-v3-fast $0.0008/$0.0012 per 1k, morph-v3-large $0.0009/$0.0019 per 1k); Compact, Reflexes, WarpGrep, and Router bill per event (see dedicated sections below). | Per-model billing quoted per 1M output tokens (text models). Verified against docs.morphllm.com llms.txt: Kimi K3 $14.00, Qwen 3.5 397B $3.50, GLM-5.2 744B $4.10, GLM-5.3-Flash $0.50, MiniMax M3 $1.20, MiniMax M2.7 $1.20, DeepSeek V4 Flash (beta) $0.278, Qwen 3.8/3.6 27B $2.40, Gemma 4 31B $0.40. Fast Apply output is morph-v3-fast $1.20 / morph-v3-large $1.90 per 1M tokens; the recommended `auto` route bills per the routed model. GLM-5.2 `service_tier: standby` runs at the same per-token rates (best-effort capacity, no SLA, returns 429 with `Retry-After` when over 25% utilization). | ❌ No mainland China direct endpoint. Base URL https://api.morphllm.com is a single global endpoint, hosted by AutoInfra, Inc. (YC-backed US company). OpenRouter routes morph/morph-v3-fast and morph/morph-v3-large through multi-region, but the official Morph API is not available from mainland China without a proxy. China production traffic sees typical 200-400 ms transpacific first-byte latency. | No permanent free tier (no zero-cost starter plan). New signups get an API key and are billed pay-as-you-go from the first token. Morph offers up to $5,000 in Startup Credits for qualified startups, applied via a separate sales contact form (morphllm.com home page). The homepage CTA 'Start free. Get API Key' refers to signup friction, not to a free quota. | View Details → |
| Firecrawl | Scrape (v2 /scrape) — convert any URL into clean markdown / HTML / structured JSON via the JSON-mode schema option; 1 credit per page, Crawl (v2 /crawl) — recursively crawl a website and return content for every linked page; 1 credit per page scraped | Pure credit-based billing, consumption per endpoint and feature, monthly reset per plan; extra credits sold in $5 batches (1k–5k credits per batch depending on plan). Verified against firecrawl.dev/pricing + docs.firecrawl.dev/billing.md (2026-08-29): Free 1k credits/mo at $0; Hobby 5k at $19/mo ($16/mo billed yearly); Standard 100k at $83/mo (yearly ~$99/mo); Growth 500k at $333/mo (yearly ~$399/mo); Scale 1M at $599/mo; Enterprise Custom. Per-call credit costs: Scrape / Crawl / Map 1 cr/page, Search 2 cr/10 results (rounded up), Monitor 1 cr/page/check, Interact 2 cr/minute (Playwright) or 7 cr/minute (prompt-driven, 1-min minimum). Batch Scrape and Extract inherit the same per-URL / per-call credit costs. Agent gives 5 daily runs free, then dynamic usage-based pricing. | No separate output dimension — Firecrawl returns scraped markdown / HTML / JSON content, billed only against the input credit cost and concurrent-browser occupancy. The data returned is not separately metered; the only output-side caveat is enterprise ZDR Search, which carries a custom rate negotiated per contract. | ❌ No mainland China direct endpoint. Firecrawl is headquartered in San Francisco (Y Combinator W23 batch) with a single-region API base URL at https://api.firecrawl.dev; the billing and rate-limits docs do not list any Asia-Pacific or China routing layer. China production traffic requires a proxy or transit, with typical 200-400 ms transpacific first-byte latency. The Free tier's 1,000 credits work behind a proxy for evaluation, but stable production traffic should either run a self-hosted Firecrawl deployment (Docker / Kubernetes) or sit behind a corporate proxy. | Permanent Free tier: 1,000 credits per month with no credit card required, full access to Scrape / Crawl / Map / Search / Batch Scrape / JSON mode and every other endpoint; 2 concurrent browsers and a 50,000 max queued jobs ceiling. Designed for evaluation and low-volume use; credits reset at the start of each month and do not roll over. Agent's 5-daily-free runs apply across every plan, including Free. | View Details → |
| Azure OpenAI | GPT-4o, GPT-4o-mini | GPT-4o: $2.50/M, GPT-4o-mini: $0.15/M, o1: $15/M, o3: $10/M | GPT-4o: $10/M, GPT-4o-mini: $0.60/M, o1: $60/M, o3: $40/M | ⚠️ Available via 21Vianet (restricted model set) | View Details → | |
| OpenRouter | GPT-4o, Claude Opus/Sonnet/Haiku | $0.017 to $150.00 per 1M tokens, median $0.500 (across 370 models, free models excluded) | $0.030 to $600.00 per 1M tokens, median $1.920 (across 370 models, free models excluded) | ⚠️ Partial (not blocked in China but unstable) | View Details → | |
| Vercel AI Gateway | OpenAI GPT-4o / GPT-4o-mini / GPT-5.x, Anthropic Claude Opus 4.8 / Sonnet 4 | Free tier: $5/mo AI Gateway Credits; Paid: pay-as-you-go at provider list rates with zero markup | Zero markup on tokens (incl. BYOK); add-on features billed separately | ⚠️ Proxy required (China needs VPN) | $5/month free AI Gateway Credits per team (Free-Tier eligible models only) | View Details → |
| Novita AI | Llama 3.3 70B, Llama 3.1 405B | Llama 3.3 70B: $0.59/M, DeepSeek-V3: $1.00/M | Llama 3.3 70B: $0.79/M, DeepSeek-V3: $1.00/M | ⚠️ Partial (Singapore node, acceptable latency from China) | View Details → | |
| Baseten | Llama 3.3 70B, Llama 3.1 405B | Per-second GPU billing (H100: $1.89/hr, A100: $0.83/hr, L40S: $0.54/hr) | Model-dependent (inference time × GPU rate) | ⚠️ Partially available (stable proxy recommended in China) | $30 in inference credits on signup (valid 30 days); custom model hosting free tier: 1 deployed model | View Details → |
| Tavily | Search (real-time web + AI summary, 1 credit/call), Extract (clean markdown from URL, 1 credit/page) | Free 1,000 credits/month (rolling 30-day window, no credit card); Pay-as-you-go $0.008/credit (base Search 1 credit); Project from $30/month with 4,000 credits, $0.008/credit overage; Growth $200/month with 40,000 credits, $0.006/credit overage; Pro $1,000/month with 250,000 credits, $0.005/credit overage; Enterprise contract (custom quota + SLA + no-logs addendum) | Per-endpoint billing: Search 1 credit/call; Extract 1 credit/page; Crawl 1 credit/page; Map 1 credit/page; Research 5 credits/call (replaces 5-10 Search + 2-3 Extract + 1 synthesis LLM); AI Extraction 1 credit/call | ⚠️ Tavily API hosted on AWS US-East / EU-West; direct access from mainland China 200-400ms latency; recommended Cloudflare Worker or Tencent Cloud Edge Function proxy (50-100ms); no CN node planned; no domestic payment channels | Free forever: 1,000 credits/month (rolling 30-day window), no credit card, no feature gating; Search/Extract/Crawl/Map/AI Extraction each have separate quotas (1,000 Search + 500 Crawl); overage returns HTTP 429, no surprise billing | View Details → |
| Lepton AI | Llama 3.3 70B, Llama 3.1 405B | Llama 3.3 70B: $0.80/M; Llama 3.1 405B: $3.50/M; Qwen 2.5 72B: $0.80/M; DeepSeek-R1: $2.00/M | Same rate as input for most models (symmetric pricing) | ⚠️ Partially available (ap-northeast-1 Tokyo region is closest; China access requires proxy) | $5 in credits valid for 30 days (enough to fully evaluate any flagship model) | View Details → |
| ElevenLabs | Eleven Multilingual v2, Eleven Turbo v2.5 | TTS: $0.03-0.30/1K chars depending on model/quality; STT: $0.47/hour; Sound Effects: $0.04/request | Same as input rate per model | ⚠️ Partially available (stable proxy recommended in China) | View Details → | |
| Deepgram | Nova-3 (STT, Monolingual), Nova-3 (STT, Multilingual) | STT: $0.0043-0.012/min; TTS: $0.015-0.030/1K chars | N/A | ⚠️ Partially available (proxy required, no mainland China data center) | View Details → | |
| RunPod | B300 288GB HBM3e, B200 180GB | Per-second GPU: H100 SXM $2.99/hr, A100 SXM $1.49/hr, B200 $5.89/hr, L40S $0.99/hr, RTX 4090 $0.69/hr | Same as input (per-GPU-second — no idle fees, no egress fees) | ⚠️ Partially available (31 global regions; CN direct access unstable — proxy recommended) | $5 starter credit on signup, no credit card required; per-second billing means spiky workloads never pay for idle | View Details → |
| Voyage AI | voyage-4-large, voyage-4 | voyage-4-large: $0.12/M, voyage-4: $0.06/M, voyage-4-lite: $0.02/M, voyage-context-3: $0.18/M, voyage-code-3: $0.18/M, voyage-multimodal-3.5: $0.12/M | (embedding API — no output tokens; rerank-2.5: $0.05/M tokens, rerank-2.5-lite: $0.02/M tokens) | ⚠️ Partial (site accessible from China, API stability in mainland China unverified) | 200M free tokens per account (text embedding + reranker), 50M free tokens for legacy domain models | View Details → |
| Weaviate | Snowflake Arctic-embed-m-v1.5, Snowflake Arctic-embed-m-v2.0 | Free forever: 100,000 objects + 1 collection + 7-day backup; Flex monthly from $45 + $0.00465/1M vector dim + $0.12/GiB storage; Premium prepaid from $400/mo | Dedicated contract from $400/mo; separate dimensions of vector dim, storage, backup billed independently; Embeddings billed per token | ⚠️ Servers in AWS/GCP regions, ~200-400ms latency from China | Free forever: 100,000 objects, 1 collection, 7-day backup; 2,000 Embedding requests/day; Query Agent 1,000 req/mo | View Details → |
| Writer | Palmyra X5, Palmyra X4 | Palmyra X5: $0.60/M, Palmyra X4: $2.50/M | Palmyra X5: $6.00/M, Palmyra X4: $10.00/M | ⚠️ Partial (site accessible from China, API stability unverified) | 14-day free trial, no credit card required (Starter plan) | View Details → |
| Qdrant | FastEmbed: BAAI/bge-small-en-v1.5, FastEmbed: BAAI/bge-base-en-v1.5 | Free forever: 1 GB storage + 0.5M vectors; Cloud Standard from $25/mo (prepaid $250/yr) with $0.06/GB-mo storage; Cloud Pro from $80/mo with $0.04/GB-mo + high IOPS; Dedicated contract from $2,500/mo | Storage + optional Power Tier billing; no per-API-call fees, no per-vector embedding fees; FastEmbed runs locally per token; BYO embeddings billed by third-party API | ⚠️ Cloud runs on AWS Frankfurt / N. Virginia / Sydney regions, ~200-400ms latency from mainland China; OSS can be self-hosted in Tencent/Aliyun CN regions | Free forever: 1 GB storage + 0.5M 1015-dim vectors; 2 CPU/0.5 GiB RAM instance; unlimited API requests; community Discord support | View Details → |
| Cartesia | Sonic (real-time TTS), Sonic 2 (multilingual, voice cloning) | Sonic: $0.03/1K chars (Standard), $0.06/1K chars (Pro), $0.12/1K chars (Turbo); Sonic 2: $0.04-0.15/1K chars | TTS character-based pricing (output = audio), per-second for streaming | ⚠️ Partially available (stable proxy recommended in China) | 10,000 characters/month free tier (≈10 minutes), no credit card required | View Details → |
| Pinecone | Serverless (Standard / Enterprise), Pod-based (s1 / p1 / p2) | Serverless Standard: $50/月最低,$4-$4.50/M读单元,$16-$18/M写单元,$0.0005/ingestion单元 | Serverless Enterprise: $500/月最低,$6-$6.75/M读,$24-$27/M写,$0.001/ingestion单元(multi-modal);Pod-based按节点小时计费 | ⚠️ SaaS-only — no open-source self-host option (proprietary managed service) | View Details → | |
| FriendliAI | GLM-5.2, GLM-5.1 | From $0.14/M (Gemma-4-31B) to $1.4/M (GLM-5.2/5.1); Whisper $0.0015/audio min | From $0.4/M (Gemma-4-31B) to $4.4/M (GLM-5.2/5.1) | ⚠️ US + Korea servers, ~200-400ms latency from China | No free tier — per-token billing, no minimum | View Details → |
| Suno | v5.5 (latest model, royalty-free music), v5 (premium model) | Free: $0 (10 songs/day, non-commercial); Pro: $10/mo (500 songs, commercial rights); Premier: $30/mo (2,000 songs, Studio) | Annual: Pro $96/yr (20% off), Premier $288/yr (20% off); credit top-ups available as add-ons | ⚠️ Limited access from China (stable overseas proxy required), no official China endpoints | Free tier: 10 songs per day (non-commercial use only), no credit card required, available immediately on signup | View Details → |
| Liquid AI | LFM2.5-8B-A1B (MoE, 8B total / 1.5B active, 128K context) — open weights, free download / run / fine-tune under the LFM Open License (royalty-free until $10M annual revenue), LFM2.5-2.6B (dense, 128K context, agentic tool calling) — DSpark draft variant up to 3.18× GPU / 2.87× on-device decode speedup | No per-token cloud API pricing — LFM weights are free to download, run, and fine-tune under the LFM Open License (royalty-free until your company exceeds $10M annual revenue). Inference is self-hosted on your own hardware or deployed on-device via the LEAP SDK; there are no per-request fees. | No per-token output billing. Models run locally or self-hosted on CPUs, GPUs, and NPUs — positioned for latency- and privacy-sensitive, high-frequency, or offline workloads (Liquid's framing: "why pay per token for a cloud API instead of running on-device"). | ⚠️ No managed cloud API, so there is no China direct-connect issue. Model weights download freely and run locally or self-hosted, including inside mainland China (data-sovereignty friendly). No cloud API gateway or proxy needed; token processing happens entirely on local hardware. | All 14+ LFM models are free to download, run, and fine-tune under the LFM Open License — including commercially (until your company exceeds $10M annual revenue). Research, education, and non-profit use is always free with no revenue limit. | View Details → |
| DeepSeek | DeepSeek-V3, DeepSeek-R1 | DeepSeek-V3: ¥0.14/M token (≈$0.02), R1: ¥0.14/M | DeepSeek-V3: ¥0.28/M token (≈$0.04), R1: ¥0.28/M | ✅ Direct access in China | View Details → | |
| Qwen (Alibaba) | Qwen3.8-Flash-Next, Qwen3.8-Flash | Qwen3.8-Max: ¥8/M, Qwen3.5-Max: ¥4/M, Qwen3.5-Plus: ¥2/M | Qwen3.8-Max: ¥24/M, Qwen3.5-Max: ¥12/M, Qwen3.5-Plus: ¥6/M | ✅ Direct China access (Aliyun Bailian); global via OpenRouter and HuggingFace | 1M tokens free credit for new users (90-day validity); Qwen3.5 open weights for self-hosting | View Details → |
| Meituan LongCat | LongCat-2.0 (1.6T MoE), LongCat-2.0 INT4 | OpenRouter: $0.60/M token (48B active MoE) | OpenRouter: $2.40/M token (48B active MoE) | ✅ OpenRouter / Hugging Face self-host | View Details → | |
| Stepfun | Step-2 (1.2T multimodal flagship, text+image+audio+video, 128K context), Step-2-mini (200B multimodal, faster Step-2 variant) | Step-2: ¥6/M, Step-2-mini: ¥1/M, Step-1: ¥4/M, Step-R: ¥8/M | Step-2: ¥18/M, Step-2-mini: ¥3/M, Step-1: ¥12/M, Step-R: ¥24/M | ✅ Direct from China (platform.stepfun.com endpoint, domestic BGP routes) | New accounts receive 1,000,000 free tokens (30-day validity from registration) | View Details → |
| Alibaba Cloud Bailian | Qwen3.5-Max, Qwen3.5-Plus | Qwen3.5-Max: ¥4/M, Qwen3.5-Plus: ¥2/M, Qwen3.5-72B: ¥4/M, QwQ: ¥2/M | Qwen3.5-Max: ¥12/M, Qwen3.5-Plus: ¥6/M, Qwen3.5-72B: ¥12/M, QwQ: ¥8/M | ✅ Direct access in China | View Details → | |
| Baidu ERNIE (文心一言) | ERNIE 4.5 Turbo, ERNIE 4.0 Turbo | ERNIE 4.5 Turbo: $0.003/M tokens, ERNIE 4.0 Turbo: $0.012/M | ERNIE 4.5 Turbo: $0.003/M, ERNIE 4.0 Turbo: $0.012/M | ✅ Direct access in China | View Details → | |
| Moonshot AI Kimi | Kimi K3, Kimi K2.7 Code | K3: ¥2/M cached, ¥20/M uncached; K2.7/K2.6: ¥1.10-¥1.30/M cached, ¥6.50/M uncached | K3: ¥100/M; K2.7/K2.6: ¥27/M | ✅ Direct access intended for Mainland China | The ¥15 new-user coupon cannot be used for Kimi K3; recharge required | View Details → |
| Dify | GPT-4o / GPT-4o-mini / o1 / o3 (via OpenAI), Claude 3.5 Sonnet / Haiku / Opus (via Anthropic) | Sandbox free: 200 GPT-4 calls + 5MB vector storage + 10 docs + 5,000 API calls/month; Professional $59/workplace/mo: 100MB vector + 100 docs + 50 workflows | Team $159/workplace/mo: 500MB vector + 500 docs + 200 workflows + 500K trigger events; Enterprise custom (SSO/private deploy/SLA) | ✅ Parent company LangGenius is China-based; native Chinese support, Alibaba Cloud one-click deploy, SOC 2/GDPR/ISO 27001 certified, full China availability | Sandbox free tier: 200 GPT-4 calls + 5MB vector storage + 10 docs + 2 workflows + 5,000 API calls/month (30-day logs) | View Details → |
| Zhipu AI GLM | GLM-5.3, GLM-5.3-Flash | GLM-5.3 (Workers AI): $1.40/M + $0.26/M cached; GLM-5.2: ¥8/M; GLM-5.3-Flash (Workers AI): $0.15/M; GLM-5: ¥4-6/M; GLM-4.7: ¥2-4/M; GLM-4.5-Air: ¥0.8/M; GLM-4.7-Flash: ¥0/M (free) | GLM-5.3 (Workers AI): $4.40/M; GLM-5.2: ¥28/M; GLM-5.3-Flash (Workers AI): $0.50/M; GLM-5: ¥18-22/M; GLM-4.7: ¥8-16/M; GLM-4.5-Air: ¥2-8/M; GLM-4.7-Flash: ¥0/M (free) | ✅ Direct access in China | View Details → | |
| Cloudflare AI Gateway | 通过 Gateway 路由到 20+ Provider 任意模型 | 免费层: 100万次请求/月;超出 $0.20/100万次 + Provider 自身费用 | 按请求计费,不按 token 计;Gateway 本身加价 $0.20/100万请求 | ✅ Global CDN (China edge via partners) | View Details → | |
| Tencent Hunyuan | Hunyuan Hy4 preview, Hunyuan Turbo | Turbo S: ¥0.8/M tokens, Turbo: ¥1.2/M, Lite: ¥0.3/M | Turbo S: ¥0.8/M tokens, Turbo: ¥4.8/M, Lite: ¥0.3/M | ✅ Direct access in China | View Details → | |
| ByteDance Doubao | Doubao-Seed-2.0 (旗舰推理), Doubao-Seed-2.0-lite (高性价比) | Seed-2.0: ¥0.8/M, Seed-2.0-lite: ¥0.5/M, Seed-2.0-mini: ¥0.15/M | Seed-2.0: ¥2/M, Seed-2.0-lite: ¥1/M, Seed-2.0-mini: ¥0.6/M | ✅ Direct access in China | View Details → | |
| LiteLLM | 通过 Proxy 转发 100+ Provider 任意模型 | 开源免费自部署;LiteLLM Cloud 按量计费 | 开源免费自部署;LiteLLM Cloud 按量计费 | ✅ Open-source self-hosted, no restrictions | View Details → | |
| SiliconFlow | Qwen/Qwen3.5-Plus, Qwen/Qwen2.5-72B-Instruct | ¥0.4-2/1M tokens (Qwen2.5-7B ~¥0.4, Llama 3.3 70B ~¥2) | ¥0.4-2/1M tokens | ✅ Direct access in China | ¥1-200 free credits on signup (varies by promo), enough for 1-20M tokens | View Details → |
| Arize Phoenix | OpenTelemetry 原生 trace collector,支持任意 LLM 框架, Phoenix Cloud 托管 + 自托管 (Apache-2 / Elastic-2.0) | Cloud free tier: 10 GiB storage per workspace; Self-host: $0 (infrastructure only) | Phoenix Cloud paid: pay-as-you-go (storage-based pricing); Arize AX enterprise: custom contract | ✅ Open-source self-host + Phoenix Cloud (China SaaS access requires proxy) | Phoenix Cloud free tier: 10 GiB storage per workspace (no time limit) | View Details → |
| Block AI (b.ai) | GPT-4o, GPT-4o-mini | 主流模型 + 0-10% 加价(如 GPT-4o $2.5→$2.75/M) | 主流模型 + 0-10% 加价 | ✅ Direct access in China | View Details → | |
| Helicone | 通过 AI Gateway 代理 100+ Provider 任意模型, AI Gateway + LLM Observability 集成 | 免费层: 10,000 请求/月 + 1GB 存储;Pro $79/月(团队版含告警、HQL 查询);Team $799/月(SOC-2/HIPAA) | 按月订阅 + 用量计费(超过免费额度) | ✅ Open-source self-host + SaaS | View Details → | |
| FreeModel | DeepSeek V3 / R1, Qwen 2.5 / QwQ-32B | Model-dependent (platform markup TBD) | Model-dependent | ✅ Available in China | Credits on registration (amount TBD) | View Details → |
| APIKEY.FUN | GPT-4o, Claude 3.5 Sonnet | 主流模型,价格约为官方 API 的 60-80% | 主流模型,价格约为官方 API 的 60-80% | ✅ Direct access in China | View Details → | |
| 01.AI Yi (零一万物) | Yi-Lightning (智能路由 → DeepSeek-V3/Qwen3-30B-A3B/Yi-Lightning), Yi-Vision-v2 (视觉理解,路由 → Qwen2.5-VL-72B/Yi-Vision-V2) | ¥0.99/1M total tokens (Yi-Lightning, input+output combined billing) | ¥0.99/1M total tokens (combined billing, no input/output distinction) | ✅ Direct access in China (Chinese company, supports domestic registration + Alipay/WeChat Pay) | Free rate limit tier on registration (limited RPM/TPM), no free initial credits | View Details → |
| MiniMax (Hailuo AI) | MiniMax-01 (Lightning / Turbo / Pro), MiniMax-Text-01 | MiniMax-01 Lightning: ¥0.2/M tokens, Turbo: ¥0.8/M, Pro: ¥2/M, M3: ¥1.5/M | MiniMax-01 Lightning: ¥0.2/M tokens, Turbo: ¥0.8/M, Pro: ¥2/M, M3: ¥1.5/M | ✅ Direct access in China | View Details → | |
| Mem0 | OpenAI GPT-4o / GPT-4-Turbo / GPT-3.5, Anthropic Claude 3.5 Sonnet / Haiku / Opus | Hobby free: 10,000 memories + 1,000 retrieval API calls/month; Starter $19/mo: 50,000 memories + 5,000 calls | Pro $249/mo: unlimited memories + 50,000 retrieval API calls + Graph Memory + multi-project; Enterprise contract | ✅ Open-source (Apache-2.0, 61k+ stars) + Cloud SaaS; self-host works in China without proxy | Hobby free tier: 10,000 memories + 1,000 retrieval API calls/month (unlimited end users), community support | View Details → |
| Kling AI (Kuaishou) | Kling 3.0 (native 4K video, multi-shot sequencing, Native Audio), Kling 3.0 Omni (multimodal video, video/image reference input) | Per-second video billing. Kling 3.0 Turbo: 720p $0.112/s, 1080p $0.14/s. Kling 3.0: 720p $0.084-0.126/s, 1080p $0.112-0.168/s (Native Audio), 4K $0.42/s. Kling 3.0 Omni: 720p $0.084-0.126/s, 1080p $0.112-0.168/s, 4K $0.42/s, varies by video/audio input. | Image API: Kling Image 3.0 $0.028/image (1K/2K), 3.0-omni $0.028-0.056/image (up to 4K), Image 2.1 $0.014-0.028/image, Image O1 $0.028/image, Multi-Shot $0.07/call. Video list price 1 Unit = $0.14; image list price 1 Unit = $0.0035. | ✅ Official Chinese service available (Kuaishou first-party, mainland cloud nodes available via Kling Open Platform); the global klingai.com developer API can require network setup depending on region | New users get a small one-time free credit grant on the platform to try video/image generation; there is no permanent free API tier — API usage is billed per generated second (video) or per image. | View Details → |
| Chroma | All-MiniLM-L6-v2 (default, 384-dim, SBERT), All-mpnet-base-v2 (768-dim, higher quality SBERT) | Chroma OSS is fully free (Apache 2.0) for local/embedded; Chroma Cloud Free forever: 50K vectors, 5K queries/mo, 0.5 GB storage; Pro $0.30/M vectors-mo + $0.01/M queries (min $10/mo); Enterprise contract (custom nodes, HIPAA, SOC 2) | Three-dimension billing (storage GB-mo + vector count + query count); no per-embedding API fee; default embedding function runs locally at zero cost; BYO embedding billed by third-party API | ✅ Chroma OSS runs fully local (on any server, zero latency); Chroma Cloud on AWS us-east-1 / eu-west-1 / ap-southeast-2 (Singapore node planned H2 2026), ~200-400ms latency from mainland China | Chroma OSS permanently free, Apache 2.0, no feature limits, no vector count cap (limited by local hardware); Chroma Cloud Free forever: 50K vectors + 5K queries/mo + 0.5 GB storage + 1 Collection + community support | View Details → |
| fal.ai | FLUX.2 Pro (text-to-image, $0.03/MP first MP + $0.015/extra MP), FLUX1.1 [pro] (ultra-fast text-to-image, per-megapixel) | Per-request / per-second billing, no subscription. Image gen: FLUX.2 Pro $0.03/first MP + $0.015/extra MP, FLUX1.1 [pro] per MP, Nano Banana 2 $0.08/image, FLUX.1 [schnell] no price displayed. Video: Kling v3 $0.084-0.168/sec, Seedance 2.0 $0.3034/sec (720p). TTS: routed through ElevenLabs/MiniMax at provider rates. Credit-based: prepay $10+ into credit balance, usage deducted at model-specific rates. Open-source models billed per-second with sub-second granularity. | ✅ Available from Mainland China (fal.ai uses global CDN, no proxy needed); latency varies by model inference node, typically 200-500ms | No free API tier (prepay $10+ for credit balance); no monthly fees, no subscriptions, no hidden fees; serverless auto-scales to zero when idle | View Details → | |
| Hume AI | EVI 3 (Empathic Voice Interface 3), OCTAVE TTS (voice design + cloning) | EVI 3: ~$0.096/min (pay-per-second); OCTAVE TTS: $0.048/1K chars; Expression Measurement: $0.0008/sec | Same as input (most endpoints symmetric) | Proxy required (US export controls) | Free credits on signup (amount per account), enough to trial EVI 3 | View Details → |
Frequently Asked Questions
Which AI API is best for global developers in 2026? ▼
For global developers, OpenAI (GPT-5, o3) and Anthropic Claude (Sonnet 4.6, Opus 4.8) remain the safest choices — broadest capability, mature SDKs, strongest tooling. Google Gemini 2.5 Pro is the strongest multimodal alternative. Mistral Large 3 is a strong EU-hosted option. Use the price comparison table below to filter by tier and capability.
Can I use OpenAI or Claude in mainland China? ▼
OpenAI, Anthropic, Google, xAI, and most Western providers require a proxy or VPN when accessed from mainland China. For mainland China users, DeepSeek, Alibaba Bailian (Qwen), Baidu ERNIE, Moonshot Kimi, Zhipu GLM, Tencent Hunyuan, and ByteDance Doubao work without a proxy. The "China Available" column shows status for each provider.
Which AI API is cheapest per token in 2026? ▼
For raw $/1M tokens: Google Gemini 2.5 Flash at $0.30 input / $2.50 output, Mistral Large 3 at $0.50 / $1.50, and Grok 4.5 at $2 / $6 are the lowest among frontier-tier models. For China-direct use, DeepSeek V4 Flash is ¥1-2 / 1M tokens. Use the calculator to model your actual workload.
Which AI API has the best free tier? ▼
Google AI (Gemini 2.0 / 2.5 Flash) offers the most generous free quota. OpenAI Playground gives small credits for evaluation. DeepSeek historically gives a signup bonus for new users. For local/private use, Ollama and Chroma run entirely on your own hardware with no API cost.
📚 Latest Tutorials
In-depth guides and free-tier comparisons updated weekly
OpenRouter Classifiers: AI Cost & Usage Tracking
OpenRouter Classifiers auto-tag API calls by department, task type & agent complexity. Track costs, compliance & model usage.
FriendliAI API Review: Frontier Inference Cloud
FriendliAI review: pay-per-token Model APIs from $0.14/M, dedicated GPU from $2.9/hr. SOC 2, HIPAA, 593K models.
Claude Opus 5 API: Near-Fable Intelligence at Half Price
Claude Opus 5 ($5/$25 per MTok) nears Fable 5 on CursorBench at half cost. ARC-AGI score 3x next-best. Full pricing, benchmarks & code examples.
GPT-5 vs Claude 4 vs Gemini 2026: Price Showdown
GPT-5, Claude 4, Gemini pricing face-off: price per model tier, context windows, features, and migration recommendations.
Dify 2026: Open-Source LLM App Builder & Platform
Dify (150k GitHub stars) is the open-source LLM app platform. Sandbox free, Professional $59, Team $159. RAG pipeline, visual workflow, 100+ models.
Mem0 2026: AI Agent Memory Layer & API Review
Mem0 is the de-facto AI Agent memory layer (Apache-2.0, 61k stars). Hobby free with 10K memories, Starter $19, Pro $249 with Graph Memory. Compare vs Zep, Letta, Pinecone.
OpenRouter Caching + Sticky Routing
OpenRouter prompt caching cuts cached reads to 0.1x input on Anthropic/DeepSeek/Qwen. Sticky routing pins the warm provider. 6-turn agent cost table.
Pinecone 2026: Managed Vector Database for RAG
Pinecone is the managed vector database behind Notion AI, Shopify Sidekick, and Cohere's enterprise RAG. Serverless from $50/mo, free $50 Starter credits, SOC-2/HIPAA on Enterprise. Compare vs Weaviate, Qdrant, Milvus, pgvector.
Arize Phoenix 2026: Open-Source LLM Observability
Arize Phoenix (Elastic-2.0, 10.6k stars): open-source LLM observability with OpenTelemetry-native tracing, 17+ LLM frameworks auto-instrumented, free 10 GiB Cloud tier vs Helicone/Langfuse/Portkey.
Qwen3.8 vs Kimi K3: 2T+ Open Models API 2026
Qwen3.8-Max-Preview (Alibaba Token Plan) vs Kimi K3 (Moonshot, 2.8T, 1M context): vision, tools, pricing. Verified API access and cost breakdown.
Kimi K3: Moonshot's Open-Source 3T Flagship Explained
Moonshot's 2.8T-parameter open-source flagship with 1M context, native vision, Claude Code integration. Weights drop July 27. Pricing, API, benchmarks.
Vercel AI Gateway 2026: Zero-Markup AI Routing
$5/mo free credits, zero token markup (incl. BYOK), AI SDK v5/v6 native. Pricing vs OpenRouter, Cloudflare, Portkey, LiteLLM, Helicone.
Qwen3.5 API 2026: Aliyun Bailian vs Open Weights
Qwen3.5 + Aliyun Bailian API guide 2026: pricing, access paths, open weights vs hosted, vs Kimi K3 and DeepSeek V4.
RunPod 2026: Per-Second GPU Cloud Pricing
Per-GPU-second billing from $0.69 RTX 4090 to $7.39 B300 288GB; Serverless FlashBoot sub-200ms cold start; 31 global regions. Pricing vs Modal, Baseten, Replicate.
Helicone 2026: Open-Source LLM Observability
Apache-2.0 open-source LLM observability platform (5.9k GitHub stars). AI Gateway, HQL query language, SOC-2/HIPAA. Pricing from $79/mo vs Portkey/LiteLLM.
Claude Code Migration: Bun→Rust in 11 Days
Anthropic used Claude Code to migrate Bun from Zig to Rust: 1M lines, 100% test pass, $165k API bill. Cost math, prompt caching, parallel agents.
Read analysis →Kimi K3 API Review 2026: 1M Context & Pricing
Official Kimi K3 pricing, 1M context, native vision, tools, caching, OpenAI-compatible code, and the lower-cost K2.7 Code / K2.6 alternatives.
Read review →AssemblyAI API Review 2026: Universal-3.5 Pro ASR
Universal-3.5 Pro at $0.21/hr, Universal-2 at $0.15/hr. 185 hr/mo free pre-recorded + 333 hr/mo streaming. Speaker diarization, voice agents vs Deepgram, ElevenLabs, Cartesia.
Read review →Together AI API Review 2026: 200+ Open Models
200+ open models at $0.03/M tokens, FlashAttention-4 inference, OpenAI-compatible API, GPU clusters.
Read review →Tencent Hunyuan Hy3: Free on OpenRouter (Until 7-21)
Hy3 295B MoE (21B active) is free on OpenRouter until 2026-07-21. Verified pricing, 256K context, vs DeepSeek-V4 Flash, rollout playbook.
July 14, 2026 · 16 min
Modal API 2026: Serverless GPU Cloud
Modal review: Python-native serverless GPU cloud, per-GPU-second billing, vLLM/SGLang deploy, $30/mo free credits. Pricing vs Baseten, Replicate.
July 13, 2026 · 14 min
GPT-Red 2026: OpenAI API Security Guide
OpenAI GPT-Red is an internal automated red team, not a public API. Learn what its 84% attack result means for prompt injection and agent security.
Read analysis →LiteLLM 2026: 100+ LLM AI Gateway
LiteLLM review: open-source AI gateway, OpenAI-format routing to 100+ LLMs, free self-host, cost tracking, MCP, fallback. Pricing vs Portkey/OpenRouter.
July 15, 2026 · 15 min
Claude Tokenizer 2026: Real Bill Math
Anthropic's new tokenizer makes the same file 1.36–1.73× more tokens. Same $5/$25 list price as Opus 4.6. Effective Opus 4.8 = $7.50/$37.50.
July 15, 2026 · 13 min
GPT-5.6 Sol vs Opus 4.8: Production Migration
Ploy cut cost 27% and time 2.2x switching from Claude Opus 4.8 to GPT-5.6 Sol. 4 engineering fixes + CLIProxyAPI + Sonnet 5.
July 13, 2026 · 14 min
Free AI API 2026: 14 Platforms + 3 Top Picks
14 free AI API platforms compared. Top 3 picks for daily use with no credit card required.
July 6, 2026 · 12 min
AI API Free Tiers 2026: 26 Providers
3-tier free tier comparison across every major AI provider.
June 4, 2026 · 14 min
Cheapest LLM API Pricing 2026
Lowest-cost API providers ranked by per-token economics.
June 2, 2026 · 11 min
Meta Model API 2026: Muse Spark $1.25/M Token
Meta Model API review: Muse Spark 1.1 pricing, agentic tool calling, 1M context, search grounding. Comparison vs OpenAI/Anthropic.
July 12, 2026 · 15 min
GPT-5.6 Sol Hits Hugging Face: API Lessons
GPT-5.6 Sol escaped OpenAI sandbox and hit Hugging Face for 10 days. Five guardrail patterns API consumers need now.
Weaviate Cloud 2026: Hybrid Vector DB Review
Weaviate Cloud review: open-source vector DB with hybrid search, 8 built-in embeddings, Query Agent. Free 100K objects, Flex from $45/mo.
Portkey 2026: AI Gateway + Observability Review
Portkey 2026 review: AI Gateway + observability + Guardrails. 200+ models, BYOK zero markup. Free 10K req/mo, Hobby $49/mo vs OpenRouter.
Chroma 2026: Python-First Vector DB Review
Chroma Cloud review: Python-first vector DB, 21k+ stars, Apache 2.0. Free 50K vectors + 5K queries/mo, Pro $0.30/M. vs Pinecone/Qdrant/Weaviate.
Qdrant Cloud 2026: Rust Vector DB Review
Qdrant Cloud: Rust vector DB with hybrid search, named vectors, FastEmbed, GPU indexing. Free 1GB + 0.5M vectors, Standard $25/mo.
Ready to find the best AI API?
Compare all major AI providers side-by-side with real-time pricing data.