FriendliAI API Review 2026: Frontier Inference Cloud
FriendliAI is a Korean-headquartered frontier inference cloud with two services: pay-per-token Model APIs for cutting-edge open-weight models and per-second Dedicated Endpoints with enterprise compliance (SOC 2 Type II, HIPAA). Its pricing page lists every model cost without a login wall. The model catalog spans 593,766 open models from Hugging Face. FriendliAI makes up for a smaller ecosystem versus Fireworks AI or Together AI with Asian-Pacific infrastructure (Seoul, Korea HQ) and a clean pricing model from a single API call to multi-GPU B300 clusters.
This review covers FriendliAI's verified pricing (July 2026), the two product surfaces, and where it wins versus the three closest US-based competitors.
TL;DR
- Model APIs from $0.14/M input (Gemma-4-31B-it) to $4.4/M output (GLM-5.2) — pay per token, no minimum
- Dedicated Endpoints from $2.9/hr (A100) to $12/hr (B300 288GB) — billed per second
- Prompt caching on 4 models — cached reads at 0.1x–0.5x of input price
- 593,766 open models deployable on Dedicated Endpoints
- Compliance: SOC 2 Type II and HIPAA
- HQ: San Francisco + Seoul — best Asian-Pacific latency
- Best for: Teams deploying frontier open-weight models in production, especially Asia-based
Why FriendliAI matters in 2026
The inference API market has consolidated around US-based providers. FriendliAI breaks the geographic monopoly: a Korean company with SOC 2 Type II and HIPAA compliance serving frontier models at competitive rates from Asian-Pacific infrastructure.
For teams deploying GLM-5.2 (Zhipu's latest flagship), DeepSeek-V3.2, or Qwen3-235B, FriendliAI's Model APIs offer these models at prices that undercut or match US peers. Its Dedicated Endpoints add a per-second GPU rental layer for predictable infrastructure without the 30–50x markup of AWS/GCP. The 593K+ model catalog enables one-click production inference for any public Hugging Face model.
FriendliAI pricing — verified 2026-07-25
All prices below were fetched live from friendli.ai/pricing on 2026-07-25. Two distinct pricing surfaces.
Model APIs — Pay per token
| Model | Input $/1M | Cached $/1M | Output $/1M |
|---|---|---|---|
| GLM-5.2 / GLM-5.1 | $1.4 | $0.26 | $4.4 |
| DeepSeek-V3.2 | $0.5 | $0.25 | $1.5 |
| Qwen3-235B-A22B-2507 | $0.2 | — | $0.8 |
| MiniMax-M2.5 | $0.3 | $0.06 | $1.2 |
| Gemma-4-31B-it | $0.14 | — | $0.4 |
| K-EXAONE-236B | $0.2 | $0.1 | $0.8 |
| Whisper-large-v3 | $0.0015 per audio minute | ||
Dedicated Endpoints — Per-second GPU rental
| GPU | VRAM | $/hr |
|---|---|---|
| A100 80GB | 80 GB | $2.9 |
| H100 80GB | 80 GB | $3.9 |
| H200 141GB | 141 GB | $4.5 |
| B200 180GB | 180 GB | $8.9 |
| B300 288GB | 288 GB | $12 |
Billing is per-second with no minimum commit. No extra charge for startup times. FriendliAI claims 2–3x higher throughput than open-source inference engines on the same GPU type.
FriendliAI vs Fireworks AI vs Together AI vs DeepInfra
| Feature | FriendliAI | Fireworks AI | Together AI | DeepInfra |
|---|---|---|---|---|
| HQ | San Francisco + Seoul | US (California) | US (California) | US |
| Compliance | SOC 2 Type II, HIPAA | SOC 2 | SOC 2 | SOC 2 |
| Asian latency | ✅ Seoul infra | ⚠️ US West only | ⚠️ US West only | ⚠️ US/EU only |
| Model catalog | 593K+ | 100+ curated | 200+ curated | 50+ curated |
| Prompt caching | ✅ 4 models | ✅ | ✅ | ✅ |
| Dedicated GPU | $2.9–$12/hr | $2–$8/hr | $2–$10/hr | Contact |
| Free tier | ❌ | $0.50 credit | $1 credit | $0.50 credit |
Pros and cons
Strengths
- Asian-Pacific infrastructure — Seoul-based hosting for teams in Korea, Japan, Southeast Asia, and Australia
- Enterprise compliance — SOC 2 Type II + HIPAA, rare in the inference API space
- Massive model catalog — 593K+ open models deployable on dedicated endpoints
- Transparent pricing — every model cost and GPU tier listed without login wall
- Per-second billing — no rounding to the hour for dedicated endpoints
- Prompt caching — cached reads at 0.1x–0.5x input cost
Limitations
- Smaller ecosystem — fewer integrations (LangChain, LlamaIndex) than US peers
- No free tier — every API call costs from day one
- Limited curated catalog — only 8 models on Model APIs; 593K require Dedicated Endpoints
- Newer entrant — smaller community, fewer tutorials
- China latency — 200–400ms from mainland China
Who should use FriendliAI?
FriendliAI is a strong choice for three groups: (1) Asia-based teams deploying frontier open-weight models who want single-digit-millisecond latency without routing through US West Coast infrastructure; (2) enterprises that require SOC 2 Type II or HIPAA compliance for inference workloads — a combination most inference API providers do not offer; and (3) teams running long-tail open models from Hugging Face that need production-grade dedicated endpoints without the markup of AWS/GCP.
It is less ideal for: (1) developers who rely on deep ecosystem integration where FriendliAI has sparser coverage; (2) teams that want a free tier for experimentation; and (3) latency-sensitive real-time applications serving mainland China users.
Our verdict
FriendliAI earns a solid recommendation. The pricing is competitive — Gemma-4-31B at $0.14/$0.4 matches Fireworks, and dedicated GPU rates ($2.9/hr A100 to $12/hr B300) are in line with US peers. What makes FriendliAI stand out is the geographic angle: teams in Korea, Japan, Singapore, and Australia can deploy frontier inference with significantly lower latency than routing through California.
The missing free tier and smaller ecosystem are real trade-offs. But for production workloads where compliance (SOC 2 + HIPAA) and Asian-Pacific performance matter more than a vast integration catalog, FriendliAI is a compelling option alongside Fireworks AI and Together AI.
Try FreeModel
One API key routes across multiple providers — DeepSeek + Qwen + Llama + OpenAI-compatible upstreams. Pairs with FriendliAI for cost-routed open-source model serving.