FriendliAI API Review 2026: Frontier Inference Cloud

FriendliAI is a Korean-headquartered frontier inference cloud with two services: pay-per-token Model APIs for cutting-edge open-weight models and per-second Dedicated Endpoints with enterprise compliance (SOC 2 Type II, HIPAA). Its pricing page lists every model cost without a login wall. The model catalog spans 593,766 open models from Hugging Face. FriendliAI makes up for a smaller ecosystem versus Fireworks AI or Together AI with Asian-Pacific infrastructure (Seoul, Korea HQ) and a clean pricing model from a single API call to multi-GPU B300 clusters.

This review covers FriendliAI's verified pricing (July 2026), the two product surfaces, and where it wins versus the three closest US-based competitors.

TL;DR

  • Model APIs from $0.14/M input (Gemma-4-31B-it) to $4.4/M output (GLM-5.2) — pay per token, no minimum
  • Dedicated Endpoints from $2.9/hr (A100) to $12/hr (B300 288GB) — billed per second
  • Prompt caching on 4 models — cached reads at 0.1x–0.5x of input price
  • 593,766 open models deployable on Dedicated Endpoints
  • Compliance: SOC 2 Type II and HIPAA
  • HQ: San Francisco + Seoul — best Asian-Pacific latency
  • Best for: Teams deploying frontier open-weight models in production, especially Asia-based

Why FriendliAI matters in 2026

The inference API market has consolidated around US-based providers. FriendliAI breaks the geographic monopoly: a Korean company with SOC 2 Type II and HIPAA compliance serving frontier models at competitive rates from Asian-Pacific infrastructure.

For teams deploying GLM-5.2 (Zhipu's latest flagship), DeepSeek-V3.2, or Qwen3-235B, FriendliAI's Model APIs offer these models at prices that undercut or match US peers. Its Dedicated Endpoints add a per-second GPU rental layer for predictable infrastructure without the 30–50x markup of AWS/GCP. The 593K+ model catalog enables one-click production inference for any public Hugging Face model.

FriendliAI pricing — verified 2026-07-25

All prices below were fetched live from friendli.ai/pricing on 2026-07-25. Two distinct pricing surfaces.

Model APIs — Pay per token

ModelInput $/1MCached $/1MOutput $/1M
GLM-5.2 / GLM-5.1$1.4$0.26$4.4
DeepSeek-V3.2$0.5$0.25$1.5
Qwen3-235B-A22B-2507$0.2$0.8
MiniMax-M2.5$0.3$0.06$1.2
Gemma-4-31B-it$0.14$0.4
K-EXAONE-236B$0.2$0.1$0.8
Whisper-large-v3$0.0015 per audio minute

Dedicated Endpoints — Per-second GPU rental

GPUVRAM$/hr
A100 80GB80 GB$2.9
H100 80GB80 GB$3.9
H200 141GB141 GB$4.5
B200 180GB180 GB$8.9
B300 288GB288 GB$12

Billing is per-second with no minimum commit. No extra charge for startup times. FriendliAI claims 2–3x higher throughput than open-source inference engines on the same GPU type.

FriendliAI vs Fireworks AI vs Together AI vs DeepInfra

FeatureFriendliAIFireworks AITogether AIDeepInfra
HQSan Francisco + SeoulUS (California)US (California)US
ComplianceSOC 2 Type II, HIPAASOC 2SOC 2SOC 2
Asian latency✅ Seoul infra⚠️ US West only⚠️ US West only⚠️ US/EU only
Model catalog593K+100+ curated200+ curated50+ curated
Prompt caching✅ 4 models
Dedicated GPU$2.9–$12/hr$2–$8/hr$2–$10/hrContact
Free tier$0.50 credit$1 credit$0.50 credit

Pros and cons

Strengths

  • Asian-Pacific infrastructure — Seoul-based hosting for teams in Korea, Japan, Southeast Asia, and Australia
  • Enterprise compliance — SOC 2 Type II + HIPAA, rare in the inference API space
  • Massive model catalog — 593K+ open models deployable on dedicated endpoints
  • Transparent pricing — every model cost and GPU tier listed without login wall
  • Per-second billing — no rounding to the hour for dedicated endpoints
  • Prompt caching — cached reads at 0.1x–0.5x input cost

Limitations

  • Smaller ecosystem — fewer integrations (LangChain, LlamaIndex) than US peers
  • No free tier — every API call costs from day one
  • Limited curated catalog — only 8 models on Model APIs; 593K require Dedicated Endpoints
  • Newer entrant — smaller community, fewer tutorials
  • China latency — 200–400ms from mainland China

Who should use FriendliAI?

FriendliAI is a strong choice for three groups: (1) Asia-based teams deploying frontier open-weight models who want single-digit-millisecond latency without routing through US West Coast infrastructure; (2) enterprises that require SOC 2 Type II or HIPAA compliance for inference workloads — a combination most inference API providers do not offer; and (3) teams running long-tail open models from Hugging Face that need production-grade dedicated endpoints without the markup of AWS/GCP.

It is less ideal for: (1) developers who rely on deep ecosystem integration where FriendliAI has sparser coverage; (2) teams that want a free tier for experimentation; and (3) latency-sensitive real-time applications serving mainland China users.

Our verdict

FriendliAI earns a solid recommendation. The pricing is competitive — Gemma-4-31B at $0.14/$0.4 matches Fireworks, and dedicated GPU rates ($2.9/hr A100 to $12/hr B300) are in line with US peers. What makes FriendliAI stand out is the geographic angle: teams in Korea, Japan, Singapore, and Australia can deploy frontier inference with significantly lower latency than routing through California.

The missing free tier and smaller ecosystem are real trade-offs. But for production workloads where compliance (SOC 2 + HIPAA) and Asian-Pacific performance matter more than a vast integration catalog, FriendliAI is a compelling option alongside Fireworks AI and Together AI.

Try FreeModel

One API key routes across multiple providers — DeepSeek + Qwen + Llama + OpenAI-compatible upstreams. Pairs with FriendliAI for cost-routed open-source model serving.

Get free credits →