Fireworks AI
Listed at https://fireworks.ai
💰 Token Pricing
| Type | Price | Note |
|---|---|---|
| Input | Serverless (Standard, $/1M tokens): DeepSeek V4 Pro $1.74、V4 Flash $0.22、Kimi K3 $3.00、K2.7 Code $0.95、MiniMax M3 $0.30、Qwen 3.7 Plus $0.40、GPT OSS 120B $0.15、GLM 5.1 $1.40。嵌入按模型档 $0.008-$0.016/1M。按需 GPU 按小时:H100 $7.00、H200 $7.00、B200 $10.00、B300 $12.00、GB300 $18.00(9 月 1 日上调)。 | per million tokens |
| Output | Serverless 输出($/1M):DeepSeek V4 Pro $3.48、V4 Flash $0.66、Kimi K3 $15.00、MiniMax M3 $1.20、Qwen 3.7 Plus $1.60、GPT OSS 120B $0.60、GLM 5.1 $4.40。缓存输入显著更低(如 DeepSeek V4 Pro $0.145)。训练:SFT 每 1M 训练 token $0.50(≤16B)到 $10.00(>300B);Serverless Training API 仅按 token 计费、无闲置 GPU 成本。 | per million tokens |
| Cache Read | 提示缓存输入 token 计费远低于全新输入——DeepSeek V4 Pro 缓存 $0.145 vs 全新 $1.74(约 -92%);适配价格见各模型档。 | Discounted |
🤖 Supported Models (9)
✨ Pros
- ✓OpenAI-compatible API — a drop-in replacement for inference and fine-tuning (same API, same SFT data format); near-zero code changes to migrate
- ✓FireAttention proprietary inference engine + quantization + speculative decoding deliver sub-second TTFT on open models — one of the fastest open-model hosting platforms
- ✓Hosts 100+ open models: DeepSeek V4 Pro/Flash, Kimi K3, Qwen 3.x, MiniMax M3, GLM 5.x, GPT-OSS, NVIDIA Nemotron
- ✓Prompt caching offers deep discounts — DeepSeek V4 Pro cached input $0.145 vs $1.74 fresh (~-92%)
- ✓Three serverless serving paths (Standard/Priority/Fast) to trade price vs latency per workload
- ✓Fine-tuning / training up to 1T+ params; Serverless Training API bills tokens only, no idle GPU cost
- ✓$1 free credit, adaptive rate limits, global multi-region deployments (GLOBAL/US/EUROPE/APAC)
⚠️ Cons
- ×No mainland China direct endpoint — production in China requires proxy/relay, 150-250ms cross-Pacific latency
- ×On-demand GPUs rise across the board on 2026-09-01 (H100 $7→$8, B200 $10→$13, GB300 $18→$20/hr)
- ×No proprietary frontier models of its own — hosts open weights (DeepSeek/Kimi/Qwen etc.) rather than R&D-ing a flagship
- ×On-demand GPUs are deployment-oriented; bursty small workloads are less flexible than per-token serverless; region-locked deployments cost 1.5x
🎯 Best For
Developers and teams that want the fastest open-model inference with an OpenAI-compatible API and no self-hosted vLLM — especially low-latency, low-cost production calls to open models like DeepSeek V4, Kimi, Qwen, and GPT-OSS.
💰 Pricing & Plans
| Service / Model | Input ($/M) | Cached ($/M) | Output ($/M) | Note |
|---|---|---|---|---|
| DeepSeek V4 Pro (Serverless, Standard) | $1.74 | $0.145 | $3.48 | V4 Pro 0813; Priority $2.61 / $0.218 / $5.22 |
| DeepSeek V4 Flash (Serverless, Standard) | $0.22 | $0.007 | $0.66 | V4 Flash 0731; connects to the Aug 2026 V4 model family |
| Kimi K3 (Serverless) | $3.00 | $0.30 | $15.00 | K3 K2.7 Code $0.95 / $0.19 / $4.00 |
| MiniMax M3 (Serverless) | $0.30 | $0.06 | $1.20 | Budget frontier-tier open model |
| Qwen 3.7 Plus (Serverless) | $0.40 | $0.08 | $1.60 | High-value mid-size open model |
| OpenAI GPT OSS 120B (Serverless) | $0.15 | $0.015 | $0.60 | GPT OSS 20B $0.07 / $0.035 / $0.30 |
| On-demand GPU — NVIDIA H100 80GB | per GPU-hr | $7.00 | — | Rises to $8.00/hr from Sep 1, 2026 |
| On-demand GPU — NVIDIA B200 180GB | per GPU-hr | $10.00 | — | Rises to $13.00/hr from Sep 1, 2026 |
| On-demand GPU — NVIDIA GB300 288GB | per GPU-hr | $18.00 | — | Rises to $20.00/hr from Sep 1, 2026 |
🔧 API & Developer Experience
- •API Style: OpenAI-compatible REST API — a drop-in replacement for OpenAI for both inference and fine-tuning (same API, same SFT data format). Switch the base URL to https://api.fireworks.ai/inference/v1 and your existing client works.
- •Serving Paths: Three serverless tiers — Standard (cost-optimized), Priority (higher throughput), and Fast (lowest latency) — so you can trade price for speed per workload without changing models.
- •Prompt Caching: Cached input tokens are billed well below fresh input (e.g. DeepSeek V4 Pro $0.145 vs $1.74 cached). Long system prompts, few-shot examples, and multi-turn contexts get large discounts automatically.
- •FireAttention Engine: Fireworks' proprietary inference engine serves quantized open models at frontier speeds — the reason GPT-OSS, DeepSeek V4, Qwen, and Llama run with sub-second TTFT on its network.
- •Model Coverage: 100+ models across text, vision, audio, image, and embeddings — including DeepSeek V4 Pro/Flash, Kimi K3/K2.7, Qwen 3.x, GPT-OSS, GLM 5.x, MiniMax M3, and NVIDIA Nemotron. Embeddings from $0.008/1M input tokens.
- •Fine-tuning & Training: Supervised (SFT), preference (DPO), and reinforcement fine-tuning for models up to 1T+ params. SFT from $0.50/1M train tokens (<16B); a Serverless Training API bills only tokens you prefill/sample/train with no idle GPU cost.
- •Free Tier & Quotas: $1 in free serverless credits to start; adaptive rate limits that grow and shrink with your usage; region-restricted single-region deployments available at a 1.5x premium.
🏎️ Fastest Open-Model Inference
Fireworks AI positions itself as the fastest platform for building with open-source models, and its core differentiator is speed on open weights. The proprietary FireAttention inference engine plus quantization and speculative decoding delivers sub-second time-to-first-token for frontier open models like GPT-OSS 120B and DeepSeek V4 Pro. Because it hosts the same models the API community actually runs — DeepSeek V4 Pro and Flash, Kimi K3, Qwen 3.x, MiniMax M3, GLM 5.x — it is a confident low-latency alternative to running vLLM yourself. DeepSeek V4 Flash pricing at $0.22 input / $0.66 output per 1M tokens makes it a strong budget pick, and the three serving paths let you tune the price-latency trade-off per workload.
🌐 Regional Availability & Latency
Fireworks runs a global fleet. Serverless inference defaults to multi-region deployments across four high-level groupings — GLOBAL, US, EUROPE, and APAC — so traffic is served close to your users and absorbs localized outages. For dedicated on-demand GPUs, single regions span 12 US sites (Arizona, California, Georgia, Illinois, Iowa, Ohio, Texas, Utah, Virginia, and three Washington zones) plus EU (Frankfurt, Iceland) and APAC (two Tokyo zones), with H100/A100/B200/H200 across them. There is no mainland China direct endpoint: as a US-based platform, direct access from the mainland requires a proxy or relay, and cross-Pacific first-byte latency typically runs 150-250ms versus ~50-80ms from a regional provider. For latency-sensitive China-facing traffic, pair Fireworks with an overseas relay or choose a regional endpoint; the Tokyo multi-region is the closest APAC option.